A voice agent that cannot be interrupted is not a conversation partner — it is a recording with a chat interface bolted on. The moment a user says “no, wait, I meant the second one” and the agent keeps talking, trust in the whole product collapses. Users do not learn to wait politely; they hang up.
Barge-in is the capability that fixes this: the user starts speaking, the agent stops speaking, and the agent responds to what the user just said. It is a small-sounding feature that touches almost every layer of a real-time voice stack — audio capture, voice activity detection (VAD), streaming speech-to-text (STT), turn segmentation, LLM generation, and text-to-speech (TTS) playback.
This guide covers the architecture behind interruption handling and the concrete steps to implement it with ZEGOCLOUD AI Agent, including the exact parameters, callback events, and tuning trade-offs you need to get right. It assumes you already have — or are about to build — a working voice agent and want barge-in to feel natural rather than glitchy.
What barge-in actually requires
Barge-in is often described as “the agent stops talking when the user speaks.” That description hides four separate problems that must all be solved at the same time:
- Detect speech, not noise. The agent must distinguish real user speech from background noise, from the agent’s own voice leaking back through the microphone, and from filler sounds like “um” that should not interrupt anything.
- Stop output immediately. Detection is worthless if the TTS pipeline keeps playing buffered audio. The current speech request has to be cancelled mid-stream, not merely marked for early termination.
- Preserve what was said. The user’s opening words arrive while the agent is still speaking, so the audio that triggered the interruption must be captured and fed into recognition. Dropping it produces an agent that reacts to the second half of every sentence.
- Re-segment the turn. Once the user is speaking, the system must decide when they have finished, so the LLM gets a complete question rather than a fragment.
These four requirements are why barge-in is a systems problem, not a toggle. ZEGOCLOUD AI Agent addresses each of them with distinct, configurable components, which is also what makes the behaviour tunable per scenario.
The interruption-handling architecture
A production voice agent handles interruption with five cooperating pieces. Each has a specific job, and each has a specific failure mode when it is misconfigured.
| Component | Job | Failure mode if wrong |
|---|---|---|
| VAD (voice activity detection) | Decides whether incoming audio is real user speech | Too sensitive: the agent interrupts itself on noise. Too strict: the user’s short “stop” is ignored. |
| Streaming STT | Transcribes the user’s audio incrementally, including audio captured during the interruption | Final transcript is missing the words spoken while the agent was talking. |
| Turn detection / segmentation | Decides when the user has finished their turn | Long questions are split into multiple LLM calls, or short pauses produce premature replies. |
| Interruptible TTS layer | Stops the current speech request when an interruption is detected | Agent finishes its sentence before listening — the classic “talking over the user” feel. |
| Proactive speech control | Lets the agent start a turn on its own | Agent only reacts; conversations stall during silence. |
VAD decides whether an interruption is real
VAD is the gate for everything else. In ZEGOCLOUD AI Agent it runs in the AI audio processing layer alongside AI noise suppression (AI ANS) and AI echo cancellation (AI AEC). AI AEC is the component that prevents the agent from interrupting itself: it removes the agent’s own voice and background music picked up by the microphone, which is a common cause of a voice agent cutting itself off mid-sentence. It also supports volume ducking and adaptive playback volume.
Sensitivity is controlled by the VAD.VoiceDetectConfig.SensitiveConfig object, documented in Voice Interruption Sensitivity Adjustment. It has three fields:
Level(integer, range 0–3) —0medium sensitivity (the default),1low,2high,3custom.MinSpeechDur(integer, milliseconds, range 0–1000) — speech shorter than this is filtered out. Only effective whenLevel = 3.EnergyThreshold(float, range 0–1) — the audio energy above which a signal counts as speech. Lower means more sensitive. Only effective whenLevel = 3.
The three presets are tuned around a real trade-off: filtering interjections and coughs versus catching meaningful short words like “hello”, “hi”, or “stop”.
| Level | EnergyThreshold | MinSpeechDur | Filters non-meaningful sounds | Catches short meaningful words |
|---|---|---|---|---|
Low (1) |
0.4 | 100 | Good | Poor |
Medium (0, default) |
0.2 | 0 | Good | Good |
High (2) |
0.1 | 0 | Poor | Good |
Start at the default. Raise to Level = 1 only in noisy environments where false interruptions dominate; drop to Level = 2 when users legitimately interrupt with very short commands. Use Level = 3 only when the presets do not fit — for example, a call-centre line with steady background chatter where you need MinSpeechDur around 100 ms and a raised EnergyThreshold.
Turn detection decides when the user is done
Interruption is only half of turn-taking; the other half is knowing when the user has finished. Because LLM input is not streamed continuously, the agent needs an explicit boundary. ZEGOCLOUD AI Agent derives that boundary from two parameters in VAD.TurnDetectConfig, documented in Speech Segmentation Control:
SilenceSegmentation(integer, milliseconds, range 200–2000, default 500) — the silence duration after which two utterances are no longer treated as one.PauseInterval(integer, milliseconds, range 200–2000) — the window within which two utterances are treated as one. Multi-sentence concatenation in ASR is only enabled whenPauseIntervalis greater thanSilenceSegmentation.
The documented best-practice configurations map directly onto product types. Short, frequent bursts (companionship) work with SilenceSegmentation = 500 and no PauseInterval. Mixed-length, latency-sensitive speech (customer service) is the recommended default: SilenceSegmentation = 500 with PauseInterval between 1000 and 1500 ms. Users who speak in longer stretches and are tolerant of latency work better with SilenceSegmentation = 1000.
This parameter is the difference between an agent that answers half a question and one that waits for the thought to land. A caller who says “I want to change my flight… to the Thursday one” should produce one LLM call, not two.
Streaming STT and the interruptible TTS layer
Two of the five components do the mechanical work of an interruption: the recogniser that captures the overlapping audio, and the speech layer that stops. Both are streaming by design, which is what makes a sub-second interruption possible.
Recognition is incremental. The client SDK delivers ASR text as it is produced, marked with start and end flags so the UI can render a partial transcript while the user is still talking, and the server receives the completed result through the ASRResult callback. Because recognition keeps running while the agent speaks whenever voice interruption is enabled, the words that triggered the interruption are part of the same recognition stream rather than a separate recording. That is how the opening words of a barge-in survive instead of being clipped.
Recognition callbacks are not identical in every interaction mode, and the difference matters when you wire UI to them. In full-duplex mode the result is sent after the user finishes speaking. In listen-only mode every result matching the segmentation rules is sent, so a long listening session can produce several callbacks. In walkie-talkie mode the result is sent after StopListening. The full comparison is in the Interaction Mode guide.
On the output side, interruption stops the current speech request rather than letting buffered audio drain: when the user barges in, the in-flight LLM request and TTS request for that round are both cancelled. If the broadcast was already partly played, the Interrupted callback’s PlayedText field tells you how far it got, which is the only reliable way to know what the user actually heard.
Where the latency budget goes
Interruption handling does not inherently add latency, but it competes for the same budget as everything else. The end-to-end loop is: user audio → VAD → ASR → LLM → TTS → playback. The interrupt path is shorter and must be faster: user audio → VAD → stop TTS. The two paths share the VAD stage, so a configuration that delays detection to avoid false positives delays the stop as well. If the interrupt path ends up slower than the ASR path, the user hears the agent continue for a moment before going quiet, which feels broken even at a few hundred milliseconds.
ZEGOCLOUD AI Agent is built for this: the product documentation describes natural voice interruption in as little as 500 ms and response latency as low as 1 second, with full streaming processing and global access through ZEGOCLOUD’s MSDN nodes. The practical implication for your own system is to keep the STT–LLM–TTS loop sub-second and to avoid introducing any synchronous work — database lookups, moderation calls, tool invocations — between ASR completion and the first TTS chunk.
How interruption is expressed in the API
ZEGOCLOUD AI Agent separates interruption into two independent mechanisms that can be combined, as documented in Interrupt Agent:
| Mechanism | Behaviour | Typical use |
|---|---|---|
| Voice interruption | The agent monitors the user’s speech while it is speaking. When the user starts talking, the agent stops the current LLM request and TTS request and begins the next round. | AI companion, voice assistant, customer service |
| Manual interruption | The agent’s current output is interrupted by an explicit API call. | Push-to-talk interfaces, noisy venues, time-limited speaking |
Voice interruption is controlled by AdvancedConfig.InterruptMode when creating an agent instance: 0 enables voice interruption, 1 disables it, and the default is 0. When voice interruption is disabled, ASR starts only after AI output (TTS playback) completes — which is exactly what you want in a push-to-talk mode where background noise would otherwise trigger constant false interruptions.
Manual interruption is a server-side call to InterruptAgentInstance using the AgentInstanceId returned by CreateAgentInstance. It can also be triggered from the client by sending a room signalling message through ZEGO Express SDK, so an in-app “stop talking” button does not need to round-trip through your own backend.
Implementation: wiring barge-in end to end
The following steps build a voice agent with working interruption handling. They assume a ZEGOCLOUD project with the AI Agent service enabled and an agent already registered — see the Voice Call quick start for the prerequisite flow, which involves creating the agent, creating an agent instance immediately after the client joins the room, and deleting the instance when the call ends.
Step 1 — Enable voice interruption and pick the interaction mode
Voice interruption is on by default, so for a natural conversation you can simply not set AdvancedConfig.InterruptMode. Set it explicitly to make the intent visible in your configuration, and set AdvancedConfig.CommunicationMode for the interaction shape you want. The three modes are documented in Interaction Mode: 0 full-duplex (default) for free conversation, 1 listen-only for transcription-only workloads, and 2 walkie-talkie for press-to-speak hardware.
{
"AdvancedConfig": {
"InterruptMode": 0,
"CommunicationMode": 0
}
}
Full-duplex and listen-only can be switched at runtime via the Update Agent Instance API, and the switch takes effect from the next user speech. Walkie-talkie mode cannot be switched to another mode.
Step 2 — Configure VAD sensitivity for your environment
Add the VAD block to the same create-instance request. Leaving it out is valid — you get Level = 0, medium sensitivity — but tuning it is the highest-leverage change for interruption quality.
{
"VAD": {
"VoiceDetectConfig": {
"SensitiveConfig": {
"Level": 1
}
}
}
}
That configuration is a good starting point for a noisy environment, since Level = 1 prioritises filtering non-meaningful sounds. For a custom profile, set Level to 3 and supply both custom fields, since MinSpeechDur and EnergyThreshold only take effect in custom mode:
{
"VAD": {
"VoiceDetectConfig": {
"SensitiveConfig": {
"Level": 3,
"MinSpeechDur": 100,
"EnergyThreshold": 0.1
}
}
}
}
Those values make detection eager on volume (EnergyThreshold = 0.1) while still filtering very short sounds (MinSpeechDur = 100). A busy open-plan office where chatter never stops needs the opposite direction: a lower Level, so quiet background voices stop registering as speech.
Step 3 — Tune turn segmentation
Set VAD.TurnDetectConfig in the same request so the agent knows when the user has finished. For a customer-service agent, the documented recommendation is SilenceSegmentation = 500 with PauseInterval between 1000 and 1500 ms.
{
"VAD": {
"TurnDetectConfig": {
"SilenceSegmentation": 500,
"PauseInterval": 1200
}
}
}
Remember that ASR multi-sentence concatenation only activates when PauseInterval is greater than SilenceSegmentation. Setting them to the same value silently disables the concatenation behaviour.
Step 4 — Subscribe to interruption callbacks
Enabling barge-in in the pipeline is not the same as knowing when it happened. To react in your application — cancelling a UI animation, discarding a pending tool call, logging the event — subscribe to the interruption callback by setting CallbackConfig.Interrupted to 1 when creating the instance, then handle the event on your callback endpoint as described in Receiving Callback.
{
"CallbackConfig": {
"ASRResult": 1,
"LLMResult": 1,
"Interrupted": 1,
"UserSpeakAction": 1,
"AgentSpeakAction": 1,
"AgentInstanceStatus": 1
}
}
The backend then delivers an Interrupted event. The Data.Reason field tells you why the interruption happened: 1 the user is speaking, 2 an LLM trigger was sent from your server, 3 a TTS trigger was sent from your server, and 4 the agent instance was interrupted explicitly through the API. That distinction matters in production: reason 1 is a natural barge-in, while reasons 2 through 4 are your own system acting on the conversation.
As of the 2026-08-21 release, the Interrupted callback also carries a PlayedText field — the text that was actually played in the current round when the broadcast was cut off. This applies both to LLM output and to speech triggered through SendAgentInstanceTTS. Use it to keep your transcript, analytics, and any text-synchronised UI consistent with what the user actually heard, rather than with what the model intended to say.
Step 5 — Correlate events with Round
Interruption produces a burst of events across two interaction chains at once, so correlating them by timestamp is fragile. ZEGOCLOUD AI Agent assigns a Round — an ascending, never-repeating identifier — to every interaction chain, and all callbacks carry it. The mechanism is documented in Round Mechanism and Callback Tracking.
A typical interruption of a broadcast in progress looks like this:
T1 AgentInstanceStatus Round 3 Speaking (AI is broadcasting)
T2 UserSpeakAction Round 4 user speech detected, new Round starts
T3 Interrupted Round 3 Round 3 interrupted
T4 AgentInstanceStatus Round 3 Idle (Round 3 terminated)
T5 ASRResult Round 4 user speech recognition result
T6 AgentInstanceStatus Round 4 Thinking (processing Round 4)
Two consequences follow directly from this model. First, when an interruption occurs the current Round terminates immediately and its status becomes Idle. Second, if the LLM is still generating when the interruption lands, the server stops generation and will not deliver an LLMResult callback for the interrupted Round. Any application logic waiting on that response will wait forever unless you handle the Interrupted event. The documented guidance is explicit: listen for Interrupted and clean up the pending state of the interrupted Round.
Step 6 — Add proactive speech without breaking barge-in
LLMs do not speak on their own, so an agent that only reacts will go silent whenever the user pauses. ZEGOCLOUD AI Agent exposes two server APIs for this, documented in Proactive Invocation of LLM or TTS: SendAgentInstanceLLM, which simulates a user message so the model generates a context-aware turn, and SendAgentInstanceTTS, which speaks a fixed string of up to 300 characters.
Both accept Priority (Low, Medium, or High; default Medium) and SamePriorityOption (ClearAndInterrupt, the default, or Enqueue with a maximum queue of 5). Those two parameters are how you decide whether a proactive utterance yields to the user:
- Welcome message that must be heard —
Priority = High. A normal-priority greeting can be cut off if the user starts speaking immediately, so high priority is what guarantees delivery. - Cold-start prompt that should yield —
Priority = Medium. The agent opens the conversation, but the user’s speech takes precedence, which is the behaviour you want in a companion or assistant product. - Critical notification that must complete —
Priority = HighwithSamePriorityOption = ClearAndInterrupt, so it interrupts whatever is playing and is not cut short. - Follow-up question after the current answer —
Priority = MediumwithSamePriorityOption = Enqueue, so it plays after the current broadcast finishes.
One detail worth knowing before you build a transcript around it: text passed to SendAgentInstanceLLM is not recorded in conversation history and is not delivered through RTC room messages, while the LLM’s response is. Text passed to SendAgentInstanceTTS is added to history by default via the AddHistory parameter.
Step 7 — Handle manual interruption and push-to-talk
Not every product wants the agent to yield to any sound. In a noisy exhibition hall or a time-limited speaking format, voice interruption produces chaos. The documented pattern is to combine settings rather than patch behaviour in the client: disable voice interruption with InterruptMode = 1, then drive turns explicitly with InterruptAgentInstance for manual interruption, and use walkie-talkie mode (CommunicationMode = 2) with StartListening and StopListening when you need press-and-hold semantics. In walkie-talkie mode a listening session typically runs within 30 seconds, and calling StopListening with CancelFlag set to true abandons the turn entirely: cached audio and recognition results are discarded, no ASR final callback is sent, no LLM or TTS request is made, and nothing is written to history.
That last behaviour is more useful than it first appears. It gives you a clean way to discard an accidental press without polluting the model’s context with a half-sentence.
Two scenarios, two tuning profiles
Customer-service agent yielding to a caller correction
A caller says: “I need to reschedule my appointment for Tuesday — actually, no, make it Wednesday afternoon.” A naive agent that segments on a short silence fires two LLM requests: one for Tuesday, one for Wednesday afternoon. The caller hears the agent confirm the wrong date before correcting itself, which is worse than saying nothing.
The fix is a segmentation profile that tolerates the self-correction, combined with VAD that catches the interjection. Set SilenceSegmentation = 500 and PauseInterval to 1000–1500 ms so the two clauses concatenate into one turn, and keep VAD at the default medium sensitivity so a normal-volume “actually, no” registers immediately. If the caller’s environment is a car or a call-centre floor, move VAD to Level = 1 so background chatter stops triggering interruptions — accepting that very short corrections may then be missed.
The callback wiring matters here too. Because the caller’s correction starts a new Round while the previous Round is still generating, the agent’s first answer will never produce an LLMResult. Handle Interrupted for Round N, discard any pending side effects keyed to Round N, and let Round N+1 proceed.
In-app assistant handling a mid-sentence question
An assistant is reading out a summary when the user asks, mid-sentence, “wait, what was the second number?” The user’s words overlap the agent’s speech, so the audio that triggers the interruption contains the question itself. This is the case where echo cancellation and audio capture behaviour decide whether barge-in works at all: without AI AEC removing the agent’s own playback from the microphone signal, the first thing the recogniser hears is the agent, and the transcript comes back mangled.
Once the interruption lands, the assistant should answer only the question asked. That means the current Round is terminated, the new Round is built from the user’s audio alone, and — because the user asked about something the agent was in the middle of saying — the system prompt should make the model aware of what was already delivered. The PlayedText field on the Interrupted callback exists for exactly this reason: it tells you what the user actually heard, so you can append that to the next turn’s context instead of re-reading the whole summary or, worse, answering as if nothing had been said.
How this compares to other agent platforms
On the platforms that document interruption explicitly, the feature set is similar — automatic interruption driven by speech detection, plus a manual interrupt path — so the difference is less about whether barge-in exists and more about how it is configured and where the control lives.
Agora’s Conversational AI Engine, documented under its Interrupt the agent mid-response guide, configures interruption through a turn_detection object on the agent-start request. Start-of-speech and end-of-speech are separate sub-objects with their own modes; its documentation states that enabling AIVAD-based interruption requires setting the end-of-speech mode to "semantic". Manual interruption is available through a REST endpoint (POST /api/conversational-ai-agent/v2/projects/{appid}/agents/{agentId}/interrupt) or through client toolkit methods on iOS, Android, and Web; the client toolkit path requires RTC SDK v4.5.1 or later and Signaling enabled for the project.
ZEGOCLOUD AI Agent splits the same problem into explicit, independently tunable knobs: AdvancedConfig.InterruptMode for whether voice interruption is on at all, VAD.VoiceDetectConfig.SensitiveConfig for how sensitive detection is, VAD.TurnDetectConfig for when a turn ends, and SendAgentInstanceLLM/SendAgentInstanceTTS priority for how proactive utterances interact with the user’s speech. Interruption control lives in the server API, and manual interruption is additionally reachable from client-side Express SDK room signalling.
The practical difference is where tuning effort goes. If your product has one conversation shape, a single well-chosen set of VAD and segmentation parameters will carry you. If you ship several — a noisy kiosk, a quiet in-app assistant, and a press-to-talk device — the value of separating sensitivity from segmentation from interruption mode is that you can ship three configuration profiles against the same integration instead of branching behaviour in your client.
Key takeaways
- Barge-in is four problems, not one. Detecting real speech, stopping TTS mid-stream, preserving the overlapping audio, and re-segmenting the turn all have to work; fixing three of them still produces a broken conversation.
- VAD sensitivity is the main quality dial.
VAD.VoiceDetectConfig.SensitiveConfig.Leveltrades filtering of interjections and noise against catching short meaningful words. The default medium level is the right starting point for most products. - Segmentation decides whether answers make sense.
SilenceSegmentationandPauseIntervalinVAD.TurnDetectConfigcontrol whether a self-correcting sentence becomes one LLM call or two. Customer-service products are documented to work best at 500 ms plus 1000–1500 ms. - Interruptions must be handled in code, not just in config. When a Round is interrupted, no
LLMResultwill arrive for it. Listen for theInterruptedcallback and clean up pending state, or your application will hang waiting on a response that was deliberately cancelled. - Latency lives in the interrupt path. Detection-to-silence has to be faster than the user’s tolerance for being talked over. Keep the STT–LLM–TTS loop sub-second and keep synchronous work out of it.
- Proactive speech is the other half of turn-taking. Priority and queueing options let a greeting be guaranteed while a cold-start prompt remains interruptible.
Frequently asked questions
What is barge-in in a voice AI agent?
Barge-in, also called interruption handling, lets a user speak while the agent is talking and have the agent stop and listen. It is what makes a voice agent feel like a conversation rather than an audio player that must be allowed to finish.
How does an agent detect that it is being interrupted?
The agent runs voice activity detection on the incoming audio while it is speaking. When speech is detected, it halts the current text-to-speech output and routes the new audio to speech recognition. Tuning VAD sensitivity — in ZEGOCLOUD AI Agent, VAD.VoiceDetectConfig.SensitiveConfig.Level — is what prevents background noise and filler sounds from triggering false interruptions.
What is proactive speech?
Proactive speech lets the agent start a turn on its own instead of only responding — greeting a caller, or reopening a conversation after silence. ZEGOCLOUD AI Agent implements it through SendAgentInstanceLLM and SendAgentInstanceTTS, both of which accept a priority so you can choose whether a given utterance can be interrupted.
Does interruption handling add latency?
Not by itself. Detection and the TTS stop happen in real time, so a correctly built pipeline interrupts without perceptible delay. What preserves the natural feel is keeping the whole speech-to-text, LLM, and text-to-speech loop sub-second; ZEGOCLOUD AI Agent is documented at interruption response times as low as 500 ms and response latency as low as 1 second.
Conclusion
Interruption handling is the feature that separates a voice agent people tolerate from one they actually talk to. The architecture is not exotic — VAD, streaming STT, turn segmentation, interruptible TTS, and proactive speech — but each piece has to be tuned against your specific conversation shape, and the interrupt path has to be handled in application code, not assumed.
ZEGOCLOUD AI Agent gives you each of those pieces as a configurable parameter rather than a black box: interruption mode, VAD sensitivity, turn segmentation, priority-controlled proactive speech, and Round-based callbacks that let you reason about what happened during an interruption. Start with the documented defaults, wire the Interrupted callback on day one, and tune sensitivity once you have real audio.
Where to go next
If you are starting from scratch, the Voice Call quick start covers agent registration and the instance lifecycle. From there, Voice AI Agent Interruption Handling walks through production interruption patterns in more depth, and ZEGOCLOUD Conversational AI is where the AI Agent service lives.
Let’s Build APP Together
Start building with real-time video, voice & chat SDK for apps today!






