Fixed silence timeouts: the endpointing setting that makes a voice agent feel like an IVR
A fixed silence threshold decides when the caller has finished. Too short splits sentences, too long adds dead air, and noise and trailing digits break it anyway.
By Voice Agent Bible Research · 5 min read
Last verified 30 Sept 2026v1.0Published 30 Sept 2026
A fixed silence timeout cannot tell a pause from an ending; it splits turns on mid-sentence pauses, holds turns open on trailing digits and noise, and adds a fixed delay to every reply.
Symptoms a buyer notices
- The agent answers half your sentence, then answers the other half, or asks you to repeat something you were still saying.
- After you read out a number and stop, the agent waits several seconds before responding.
- On a noisy line or with background noise the agent seems not to notice you finished.
- Every reply has the same fixed pause before it, so the conversation has an IVR rhythm even when the answers are good.
How it shows up
You are giving the agent your date of birth. You say the day and the month, pause for a second to remember whether you said the right year last time, and the agent replies to the day and the month. You finish the year into its reply. It stops, apologises, and asks again. Later you read out a phone number, digit by digit, and stop. Nothing happens. Four seconds, five, then the agent confirms the number. Both failures come from the same setting.
The setting is a silence timeout: the agent decides you have finished speaking when the line has been quiet for a fixed length of time, commonly somewhere between half a second and a second and a half. Configure it short and the agent jumps into every pause. Configure it long and every reply carries that delay, and the conversation acquires the rhythm of a menu system: you speak, you wait, it speaks. The third symptom is subtler. On a line with comfort noise, a fan in the room or a television, the quiet never arrives, and the agent waits much longer than the configured value because the detector never saw silence.
Why buyers care
Turn-taking is the first thing a caller feels, before accuracy and before the quality of the answers. An agent that interrupts reads as rude. An agent that leaves long gaps reads as not listening. A widely repeated rule of thumb in buyer guides puts the boundary near 800 milliseconds for a conversational feel and about 1.2 seconds for an IVR feel; whatever the exact numbers, a fixed timeout at the long end of that range is paid on every turn, before the language model or the voice have done any work.
No single value is right for a whole call. A short yes wants a fast endpoint. A read-out account number, spoken in groups with pauses, wants a slow one. A caller thinking aloud wants patience; a caller giving a crisp instruction wants speed. A fixed timeout is a compromise that is wrong for some part of every conversation, and the parts it is wrong for (identity checks, numbers, hesitant callers) are the parts that matter most.
The mechanism
A fixed endpointer measures elapsed quiet since the last detected speech and fires when it exceeds a threshold. It knows nothing about what was said. Three consequences follow, and the research team's builds hit all three while producing recorded caller audio.
Punctuation plus pause splits the turn. A line written as two sentences, with a period and a natural pause between them, was split into two turns by a fixed endpointer, and also by a streaming end-of-turn model that treated the full stop as a strong ending cue. The fix on the recording side was to write one sentence per line and join clauses with commas or "and". Real callers do not follow that rule, so the failure moves from the recording studio to the phone line.
Trailing digits hold the turn open. A line ending in a bare digit string ("...seven, seven, zero, three, one.") was finalised 5 to 7 seconds late, three takes in a row, on a fixed-endpointing recogniser. The recogniser's number formatting waits for more digits that might continue the number; a trailing word ("...one, for our patient") released it. Callers reading a phone number stop on a digit almost every time.
Noise floors defeat silence detection. Recorded caller audio with a synthetic noise floor of a few least-significant bits, far below anything audible, held the turn open and produced finals about five seconds late on a fixed-endpointing recogniser. Replacing the gaps with true digital silence fixed it. A live phone line never carries true digital silence: the carrier fills quiet with comfort noise. The detector therefore rarely sees the condition it is waiting for and falls back to whatever longer timeout the platform imposes.
Streaming end-of-turn models take a different approach: they predict from the words and the prosody whether the speaker has finished, and a timeout is the fallback rather than the rule. They are not immune (the period-plus-pause split appeared there too), but they do not need digital silence and they do not treat a trailing digit as an open number.
Evidence
From the research team's demo builds, method on the methodology page. Same language model, same tools, same recorded callers; only the recogniser and its turn-detection changed.
| Turn detection | Voice-to-voice median | Notes |
|---|---|---|
| Streaming end-of-turn model with custom vocabulary | 1.5 to 2.0 s | Native end-of-turn and barge-in; the smoother default for live plays |
| Fixed-endpointing recogniser, 700 ms timeout, custom vocabulary | 2.3 to 3.2 s | Slower first turn; late finals on trailing digits and on noise floors |
- Trailing digit strings: finals 5 to 7 seconds late on the fixed-endpointing configuration, three takes in a row; resolved by ending the line on a word.
- Synthetic noise floor of a few least-significant bits between lines: finals about 5 seconds late on the fixed-endpointing configuration; resolved by digital silence, which a live line cannot provide.
- Period plus pause mid-line: split turns on both configurations; resolved on the recording side by one sentence per line.
- In an adjacent live-microphone setup, the agent's own voice reaching the microphone through the speakers triggered the fixed-endpointing recogniser's speech-start and cut the agent off. Gating barge-in on local microphone energy (above roughly minus 34 dBFS in the last 350 ms) fixed it; the streaming model was less prone but not immune.
The 0.8-second difference in median voice-to-voice is paid on every turn of every call, and it is entirely turn-detection: the model and the voice were unchanged.
How to test for it in a demo
Every probe here works over an ordinary phone call. Run each three times.
- The mid-sentence pause. "I'd like to book for Thursday" (pause about one second) "or Friday if Thursday is full." Pass: one reply covering both days. Fail: a reply to Thursday alone, then confusion.
- The trailing number. Read a ten-digit phone number, digit by digit, in groups of three, and stop on the last digit. Time the gap to the agent's first word. Pass: about a second. Fail: three seconds or more, or the agent asks whether you are still there.
- The noisy room. Repeat probe 2 with a fan or a television near the handset. Pass: the same gap as before. Fail: the gap grows, or the agent never responds until you speak again.
- The hesitant caller. Give your name as "it's, um, (pause) Alexandra, with an x." Pass: the agent waits and captures the spelling. Fail: it asks for your name again before you finish.
- The fast yes. Answer a yes-or-no question with a crisp "yes". Time the reply. If probe 2 and probe 5 have the same gap, the endpoint is a fixed timer and is being paid on every turn.
- Ask the question. What decides that the caller has finished? A fixed duration is the anti-pattern. A content-aware model with a bounded fallback is the good answer, and the vendor should be able to show the turn-split and late-final rates they measured.
Questions to ask vendors
- How does the agent decide the caller has finished speaking, and is that decision content-aware or a fixed silence duration? A good answer names a model that uses what was said and how, with a bounded fallback, and quotes measured split and late-final rates.
- What is your median voice-to-voice latency on tool-backed turns, and how much of it is waiting for the endpoint? A good answer breaks the end-of-turn component out, from real calls.
- What happens when a caller reads out a number and stops, on a line with comfort noise? A good answer is that the turn closes in about a second without needing digital silence, and that trailing digits and noisy lines were tested specifically.
- If the endpoint is a timeout, can it be changed per step of the conversation? A per-step value (patient during number capture, fast on confirmations) is a partial mitigation and a sign the vendor understands the trade-off.
Questions to ask vendors
- 01
How does the agent decide the caller has finished speaking, and is that decision content-aware or a fixed silence duration?
A good answer: A model that predicts end of turn from what was said and how it was said, with a bounded fallback timeout, and measured turn-split and late-final rates on real calls.
- 02
What is your median voice-to-voice latency on tool-backed turns, and how much of it is waiting for the endpoint?
A good answer: A measured number for tool turns with the end-of-turn component broken out, from real calls rather than a lab.
- 03
What happens when a caller reads out a number and stops, on a line with comfort noise?
A good answer: The turn closes within about a second because the detector does not need digital silence; the vendor has tested trailing digits and noisy lines specifically.
Related
- Best practice
- Best practice
- Best practice
- Use case