Skip to content
Voice AgentBible

Barge-in done right: gate interruption on microphone energy and know your echo path

Interrupting the agent must need real caller energy at the microphone, not just a speech-start event; the agent's own voice and line noise cause false barge-ins.

By · 5 min read

Last verified 30 Sept 2026v1.0Published 30 Sept 2026

Best practiceTurn-takingimpact medium

Barge-in must require real caller energy at the microphone over a recent window, not just a speech-start event; the agent's own voice and line noise both trigger false interruptions.

Why buyers care

Barge-in is the ability to interrupt the agent and have it stop. Callers expect it because people do it. Without it, every correction waits for the agent to finish a sentence it should not have started. With it done badly, the agent interrupts itself.

The second failure is the one presenters report as "the voice keeps breaking". The mechanism is mundane: the agent's speech leaves the speakers, enters the microphone, the speech model fires a speech-started event, and the client cuts the agent's audio to let the "caller" speak. Nobody spoke. In the research team's demo builds this appeared the first time a demo ran on a laptop's speakers instead of headphones, and it was reported as a defect in the voice.

Buyers care because the echo path is different for every channel their callers use, and a demo on headphones in a quiet room exercises none of them.

The mechanism

A barge-in has three parts: something says the caller started speaking, the client stops playing agent audio, and the pipeline decides what to do with the sentence the agent did not finish.

The trigger is the weak part. Most speech models emit a speech-started event from a voice-activity detector tuned to be sensitive, because missing a real onset is worse for transcription than a false start. Sensitive is wrong for barge-in. Four things trigger it falsely.

False triggerWhere it comes fromWhat the caller experiences
Acoustic echoThe agent's voice through speakers into the microphoneThe agent stops itself mid-sentence
Comfort noiseTelephony gateways synthesise a noise floor during silenceSpeech-started fires on nothing; end-of-turn is held open
Room and line noiseCoughs, keyboards, traffic, a second personThe agent stops for sounds that are not speech
Gate chatterCarrier or app-side gates that hard-mute quiet audio and reopen abruptlyClicks that look like onsets; quiet first syllables lost

Gate the trigger on local energy. In the research team's builds, live-microphone barge-in was gated on the microphone's own signal: interrupt only when the root-mean-square level over the last 350 ms exceeded roughly minus 34 dBFS. Echo through speakers and the residual after echo cancellation sit below that in a normal room; a caller's voice sits above it. The threshold is a starting point, not a constant: tune it per channel with a recording of each.

Or gate on a transcribed word. The stricter alternative is to interrupt only when the speech model has produced an actual word, not just a speech-started event. It adds a few hundred milliseconds before the agent stops, and it removes almost every false trigger. For agents in noisy environments this is often the right trade.

Use echo cancellation, and do not rely on it. Browsers and handsets cancel echo, carriers cancel echo, and a residue still gets through, especially at the start of the agent's sentences and on cheap speakers. Treat cancellation as reducing the residue, and the energy gate as the decision.

Know what telephony does to onsets. Some gateways and softphones hard-gate quiet audio to save bandwidth. When the caller starts speaking softly, the first tens of milliseconds are cut before the gate opens. Short answers ("yes", "no", a single digit) lose their onset and are mis-heard or missed. Streaming end-of-turn models are less prone to holding a turn open on comfort noise than fixed-silence endpointing, but none are immune to a clipped onset. Meanwhile a synthetic noise floor between utterances can hold a fixed-silence endpointer open for seconds. See ASR hallucination on silence for the related failure where the model invents words in the noise.

Decide what happens to the unfinished sentence. When the agent is cut off, the language model still believes it said the whole thing. Good pipelines feed back what was actually heard, so the agent does not later refer to an offer the caller never heard.

Evidence

The research team's demo builds in 2026 supply the pattern. Method: the same agent was run in a real browser with a live microphone through laptop speakers and through headphones, and separately with scripted caller audio baked with digital silence and with a low-level synthetic noise floor between lines; interruptions, held turns and lost onsets were read from the event timeline and the recordings.

  • With a fixed-endpointing speech model and no energy gate, live-microphone runs on speakers showed the agent's audio stopping mid-sentence with no caller speech. The trigger was the agent's own voice. Gating barge-in on microphone energy above roughly minus 34 dBFS over the last 350 ms removed the false stops in subsequent runs; a streaming end-of-turn model was less prone but still triggered occasionally without the gate.
  • Scripted caller lines followed by a synthetic noise floor of a few least-significant bits held a fixed-silence endpointer open, and final transcripts arrived about five seconds late. The same lines followed by digital silence finalised promptly. A streaming end-of-turn model tolerated the noise floor.
  • On carrier-gated audio reviewed in earlier investigations, hard gates reopened abruptly and pre-clipped quiet onsets, so short answers were lost or mis-heard even though the rest of the call transcribed well.
  • The agent's greeting audio arrived within about a second but played for eight to twelve seconds. A test rig that started the caller on arrival rather than at the end of playback produced a caller talking over the greeting on every run. The same mistake in a real client is a barge-in with nobody speaking.

How to test for it in a demo

  1. Speakers, not headphones. Ask for the demo to run on the laptop's speakers at normal volume. Listen for the agent stopping mid-sentence with nobody speaking. If the vendor will not, ask why.
  2. Interrupt on purpose. While the agent is mid-offer, say "actually, make it Tuesday". Time how long its audio takes to stop. A few hundred milliseconds feels natural; over a second feels like it did not hear you. Check that its next sentence acknowledges the interruption rather than finishing the old thought.
  3. Interrupt by accident. Cough. Tap the table. Have someone speak across the room. The agent should keep talking.
  4. Say something short and quiet. Answer a question with a soft "yes" or a single digit. Then do the same through the vendor's phone number from a mobile. Compare what was heard.
  5. Change channel. Browser, softphone, mobile through the carrier. Barge-in that works on one and not another is a channel problem the vendor should already know about.
  6. Ask what was heard. After an interruption, ask the agent to repeat what it offered. If it repeats the whole sentence including the part you never heard, the pipeline does not track cut-offs.

Record every call; the recording is the evidence. A harness that never plays audio into a microphone cannot show any of this, which is why harness-only testing misses it.

Questions to ask vendors

The three frontmatter questions: what triggers a barge-in and how the agent's own voice is excluded, how interruption behaves on a carrier line with comfort noise compared with a browser, and how fast the audio stops and what happens to the unfinished sentence. Vendors who have shipped on phone lines will answer per channel. Vendors who have shipped on headphones will answer once.

Questions to ask vendors

  1. 01

    What exactly triggers a barge-in on your platform, and how do you stop the agent's own voice from triggering it?

    A good answer: A speech-start event gated on local microphone energy over a recent window, or on a transcribed word, with echo cancellation described for each channel the platform supports.

  2. 02

    How does interruption behave on a carrier phone line with comfort noise, compared with a browser?

    A good answer: A specific answer per channel, with the noise handling stated, and an admission that gateways which gate quiet audio can clip onsets.

  3. 03

    How quickly does the agent's audio stop after a caller starts speaking, and is what was left unsaid dropped or resumed?

    A good answer: A measured stop time in the low hundreds of milliseconds, a deliberate policy on resumption, and the transcript showing what the caller actually heard.