Stall watchdogs: detect a silent agent or a lost transcript and recover once
Hosted agents go silent. A watchdog that arms only when the transcript is missing, clears on any agent activity and replays audio once recovers without duplicates.
By Voice Agent Bible Research · 5 min read
Last verified 30 Sept 2026v1.0Published 30 Sept 2026
A hosted voice agent can go silent for a minute or more mid-call; a well-armed watchdog detects it in seconds and recovers exactly once, without repeating any action the caller already triggered.
Why buyers care
Somewhere in a hosted voice pipeline there is a queue, and sometimes the queue stops. The caller finishes a sentence and nothing happens: no acknowledgement, no tool call, no audio. In the research team's demo builds these silences lasted 60 to 120 s and then released everything at once in a burst. The caller had hung up long before.
Buyers care for two reasons. First, the stall is intermittent and lives in the hosted parts of the stack, so it never appears in a vendor-run demo and always appears in the first week of production. Second, the obvious fix is dangerous. Reconnecting and replaying the conversation repeats the tool calls the caller already triggered: a second booking, a second charge, a second dispute. The right design detects the stall in seconds, recovers exactly once, and proves it did not duplicate anything.
The mechanism
Two different stalls need two different watchdogs.
| Stall type | What the caller did | What arrived | What never arrived | Watchdog arms on |
|---|---|---|---|---|
| Reply stall | Spoke a full turn | The transcript | Any thinking, tool request or audio from the agent | The caller's transcript, if no agent activity follows |
| Transcript stall | Spoke a full turn | Perhaps a speech-started event | The transcript itself | The end of a loud utterance detected from the audio, if no transcript has already arrived |
The stall signature is worth knowing because it tells you where the problem is. A burst release has identical timestamps on every event, turn latencies under 200 ms that are physically impossible, and a tool result minutes after the transcript that triggered it. If a vendor's log shows that shape, the platform stalled, not the application.
Arming rules. A watchdog that waits for a transcript must arm only if that transcript has not already arrived (compare the last transcript time with the utterance start). It must clear on any agent activity: a tool request, the agent starting to speak. It must never clear on "thinking", because thinking is the phase that hangs. It must re-arm after the application returns a tool result, because the model can hang on the second call of a tool turn as easily as the first. And a spoken filler is not a reply; a filler must not disarm the reply watchdog.
Detecting the transcript stall. The application watches the caller's raw audio with a simple energy gate (in the research team's builds, energy above roughly minus 39 dBFS, an utterance being at least 0.6 s loud and ending after 0.7 s quiet), keeps a ring buffer of the last 45 s of audio, and arms a 7 s timer at utterance end. If the transcript lands first, the timer never arms.
Recovering. Reconnect with the conversation history. Then, and this was probe-verified in the research team's builds, do not expect the history alone to answer the caller's trailing turn: it will not. Remove the unanswered caller turn from the history sent at reconnect and inject it as a fresh caller message once the new session is ready. For a transcript stall, replay the buffered utterance into the new session instead, with about 320 ms of pre-roll, at exactly real time in small frames, while the live microphone is muted.
Bounding it. Cap automatic recoveries per session (four in the research team's builds), tell the caller or presenter what happened in one plain sentence, and rely on duplicate suppression on every consequential tool so a repeated request returns the original result.
Knowing a false alarm. A stall entry whose caller text is about one watchdog interval old is not a stall. It is the watchdog arming after the transcript already arrived. Before shipping a watchdog, run a full scripted call and read the stall log against the event timeline.
Evidence
The research team's demo builds in September 2026 supply the pattern. Method: scripted callers in a real browser against a hosted voice-agent platform; every event was logged with a timestamp; stalls were counted by grouping the caller text that preceded each one and whether a tool request followed; recoveries were probed in isolation with small scripts before being trusted in a full run.
- Reply stalls of 60 to 120 s followed by a burst release were observed on tool-call turns and not on plain turns, with the platform emitting a warning that it had waited for a response. The same warning appeared in a control demo on the same day, attributing the stall to the hosted path rather than to either application.
- Reconnecting with history alone did not answer the trailing caller turn. Popping the turn from history and re-injecting it did: reconnect in about 0.9 s, correct answer, earlier tool results carried over.
- The first transcript-stall recovery interleaved live silence with the replayed audio. The utterance was chopped into fragments and produced seven duplicate tool calls. Muting live audio during the replay and pacing at exactly real time fixed it; pacing at twice real time also degraded the transcript.
- A transcript watchdog armed after a 0.7 s quiet detector fired false stalls on most tool turns, because the speech model returned the transcript faster than the detector fired. The recovery then produced the overlaps and cut-offs it was meant to prevent. Arming only when the transcript had not already arrived removed the false alarms.
- In one real-browser run the reply watchdog fired three times and the call still completed every step with zero overlapping audio. The re-asked turn appeared as a second caller bubble, which is the expected trace of a recovery.
How to test for it in a demo
Stalls cannot be summoned on demand, so the demo test is partly documentary.
- Ask for a stall log. Every production platform has sessions where the agent went silent. Ask to see one with timestamps: when the caller finished, when the watchdog fired, when audio resumed, what the caller heard.
- Ask about the caller's experience. What does a caller hear at five seconds of silence? At ten? Is there a spoken cue, a hold tone, a transfer? "Nothing, it just waits" is an answer.
- Ask about the bound. How many automatic recoveries before the call is transferred or closed gracefully? An unbounded loop is a caller trapped in a reconnecting agent.
- Check duplicates after any hiccup. If anything odd happens in the demo (a pause, a repeated sentence), look at the sandbox system of record afterwards. One write per action.
- Ask how false alarms were found. A team that has tuned a watchdog will describe reading the stall log against the timeline and finding entries one interval old. A team that has not will describe the timeout value.
- Ask what clears the watchdog. If "the agent started thinking" is on the list, the watchdog does not work, because thinking is the phase that hangs.
None of this appears in a harness that does not play audio or wait real time; see harness-only testing.
Questions to ask vendors
The three frontmatter questions: what the caller experiences when a model stops responding, how a recovery avoids repeating a tool call, and how false alarms are distinguished from real stalls. A vendor who has lived through this will answer with a log. A vendor who has not will tell you it does not happen on their platform, which is the answer to write down and revisit in week two.
Questions to ask vendors
- 01
What happens to the caller when the language model or the speech model stops responding mid-call, and how quickly do you detect it?
A good answer: A watchdog measured in single-digit seconds, a spoken or visible cue to the caller, a bounded number of automatic recoveries, then a transfer or a graceful close.
- 02
How does a recovery avoid repeating a tool call the caller already triggered?
A good answer: Duplicate suppression on consequential tools, history carried into the new session, and the unanswered caller turn re-asked once rather than the whole exchange replayed.
- 03
How do you know a watchdog firing was a real stall and not a false alarm?
A good answer: A stall log compared against the event timeline; a stall whose transcript is about one watchdog interval old is a false alarm, and the watchdog is tuned until those disappear.
Related
- Anti-pattern
- Anti-pattern
- Anti-pattern