Skip to content
Voice AgentBible

Fillers that talk over the caller, and silently disarm your stall detection

Injected acknowledgements look like a latency fix. In a real browser they were rejected, spoke over the caller, or fooled the logic that marks a reply done.

By · 5 min read

Last verified 30 Sept 2026v1.0Published 30 Sept 2026

Anti-patternTurn-takingseverity medium

An injected filler is a second speech stream; it collides with the caller, confuses reply completion, and if it counts as a reply it hides the stall it was meant to cover.

Symptoms a buyer notices

  • The agent says 'one moment' while the caller is still talking, or right on top of the caller's last word.
  • After a filler the agent goes quiet for a long time and nothing recovers the call; the filler was counted as the answer.
  • The next caller line or the next step fires early because the filler's audio-done event was taken as the end of the reply.
  • Filler phrasing repeats identically on every tool turn, which callers notice within two turns.

How it shows up

The agent needs to look something up. To hide the wait, it says "let me check that for you" and then, a second or two later, gives the answer. That is the intent. What you hear in a real call is different. The caller is still finishing their sentence and the filler starts on top of them. Or the filler never plays at all, because the platform refused to inject speech while the caller's audio was active, and the wait is as silent as before. Or the filler plays, the agent goes quiet, and nothing happens for a long time, because the system decided the filler was the reply and stopped waiting for one.

A second-order symptom appears in scripted or automated tests. The step after a tool turn fires early. The client logic that decides "the agent has finished speaking" saw an audio-complete event from the filler, treated it as the end of the reply, and moved on while the real answer was still being generated. The caller's next line then lands on top of the tool-backed sentence, and the transcript shows two people talking at once.

Why buyers care

Fillers are sold as a latency fix and they are usually a symptom of one. A tool turn on a typical managed stack is two language-model calls plus the tool itself; in the research team's builds each model call took roughly 0.7 seconds and each additional chained tool added about a second. If the wait is long enough to need covering, the first question is why it is that long.

The costs of getting fillers wrong are asymmetric. A well-timed filler saves the caller a second of uncertainty. A badly timed one interrupts the caller, who stops, restarts, and produces a split turn that the agent now has to reassemble. A filler that counts as a reply can hide a real stall. In the research team's builds a managed agent occasionally went silent for 60 to 120 seconds after a caller turn and then released everything in a burst; a stall watchdog that had been cleared by a filler would have let the caller sit in that silence.

The mechanism

A filler is an instruction to the platform to speak a sentence the model did not generate as part of its reply. That creates a second speech stream alongside the real one, and three things go wrong at the seams.

The caller's audio state. Voice-agent platforms gate agent speech on whether the caller is speaking. An injected filler arrives at whatever moment the tool call started, which is often while the end-of-turn detector is still deciding whether the caller has finished. In the research team's real-browser runs, injected acknowledgements were either rejected by the platform because the user was still speaking, or accepted and spoken over the caller. Neither is the intended behaviour, and which one you get depends on timing you do not control.

Reply completion. Every client, whether a browser page, a telephony bridge or a test harness, needs a rule for "the agent has finished this reply". The natural signals are an audio-complete event and a quiet window. A filler emits its own audio-complete event. A client that waits for that event, or for the quiet that follows the filler, concludes the reply is done and proceeds, while the tool-backed sentence is still on its way. The fix the research team's builds settled on is that a reply is complete only when the agent has produced spoken text after its last tool event, that speech's audio-complete has arrived, everything queued has played, and no new audio has arrived for about 1.2 seconds. A filler before the tool result does not count.

Stall detection. A stall watchdog arms on every caller turn and is cleared by agent activity. If "agent activity" includes fillers, the watchdog is disarmed the moment the filler plays, and a subsequent stall in the think or tool path goes undetected. The rule that survived testing is that a filler is not a reply: the watchdog is cleared by a real reply or a tool event, is never cleared by the thinking phase (which is exactly the phase that hangs), and is re-armed after every tool response is sent back to the model.

There is also a build-dependence trap. On an earlier build of the same stack, fillers measurably improved perceived latency in a scripted harness. On a later build, in a real browser, they were rejected or overlapped. A filler policy proven on one version is not proven on the next; it has to be re-measured after every platform change.

Evidence

From the research team's demo builds, method on the methodology page:

  • Voice-to-voice median on tool-backed turns was 1.5 to 2.0 seconds with a streaming end-of-turn recogniser and 2.3 to 3.2 seconds with a fixed-endpointing recogniser, same language model, fillers off. Reducing the number of chained tool calls per turn moved a two-tool turn from about 2.3 to 2.5 seconds into the same class as a one-tool turn. That is where the latency work belongs.
  • With fillers enabled on the later build, injected acknowledgements in the real-browser gate were rejected with a "user is currently speaking" reason or were spoken over the caller; their extra audio-complete events caused the reply-complete logic to release the next scripted caller line early.
  • Waiting only for an audio-complete event or only for a quiet window, rather than for the full reply-complete rule, produced 409 overlapping audio chunks on four of eight caller lines in one real-browser run.
  • In a separate incident a managed agent's pending think never delivered a real reply after a spoken filler; the filler played and the turn was lost. The stall watchdog, armed on the caller turn and not cleared by the filler, was what recovered the call.
  • After the measurements, the shipped configuration ran fillers off, with the tool wait covered by a faster end-of-turn detector and fewer tools per turn. Fillers were kept as an option only for waits over about 1.5 seconds, and only with the "filler is not a reply" rule enforced.

How to test for it in a demo

  1. Talk through the filler. Ask a question that needs a lookup and then keep talking: "and also, is parking available?" Pass: the agent either waits for you to finish or is cleanly interrupted, and the real answer covers both questions. Fail: "one moment" lands on top of you, or the second question is lost.
  2. Start early. As soon as a filler ends, begin your next sentence. Pass: the tool-backed answer still arrives and the agent handles the overlap gracefully. Fail: the answer is dropped or plays over you.
  3. Count fillers. Over ten tool turns, note whether the phrasing repeats identically. Identical phrasing is a sign of a fixed injected string rather than a generated reply.
  4. Force a stall. Ask the vendor to simulate a slow tool (a five-second lookup). Pass: the agent covers the wait once, then either answers or explains the delay and recovers; if the tool never returns, the agent recovers within a stated time. Fail: a filler, then silence.
  5. Ask for the numbers without fillers. A vendor who can only quote latency with fillers on has not measured the underlying turn.

Questions to ask vendors

  • When the agent says "one moment", what decides whether that sentence is allowed to start, and what happens if the caller is speaking? A good answer includes gating on caller speech, barge-in on the filler, and no orphaned completion events.
  • Does a filler count as a reply for your stall detection? The only good answer is no, with a description of what does clear the watchdog and when it is re-armed.
  • What is your median voice-to-voice latency on tool-backed turns without fillers, and what did you do to reduce it before adding them? A good answer leads with tool-count reduction and end-of-turn detection, then treats fillers as a last resort.
  • When did you last re-measure filler behaviour after a platform update? A good answer is a date and a method; fillers proven on a previous build are unproven on this one.

Questions to ask vendors

  1. 01

    When the agent says 'one moment', what decides whether that sentence is allowed to start, and what happens if the caller is speaking?

    A good answer: It is gated on the caller not speaking, it is barge-in interruptible, and if rejected the real reply still arrives without any orphaned audio events.

  2. 02

    Does a filler count as a reply for your stall detection?

    A good answer: No. The stall watchdog is cleared only by a real reply or a tool event, and re-armed after every tool response.

  3. 03

    What is your median voice-to-voice latency on tool-backed turns without fillers, and what did you do to reduce it before adding them?

    A good answer: A measured number for tool turns specifically, with the tool-count and end-of-turn work described before fillers are discussed.