Skip to content
Voice AgentBible

The latency budget: why about 800 ms voice-to-voice is the line

Voice-to-voice latency is end-of-turn detection plus one or two model calls plus first audio. Where the budget goes, and how to time it yourself in a demo.

By · 4 min read

Last verified 30 Sept 2026v1.0Published 30 Sept 2026

Best practiceLatencyimpact high

Voice-to-voice latency is end-of-turn detection plus one or two model calls plus time to first audio; under about 800 ms feels conversational, over about 1.2 s feels like an IVR.

Why buyers care

Latency is the one thing every caller notices and no spec sheet shows honestly. Vendors quote a single number, usually for the greeting or for a turn with no tool call. Callers experience the slowest turns: the ones where the agent checks availability, looks up an account or writes a booking.

Two rule-of-thumb thresholds are repeated across buyer guides. Under about 800 ms from the end of caller speech to the first agent audio feels conversational. Over about 1.2 s starts to feel like an IVR, and callers begin talking over the agent or repeating themselves. Vendor pages draw the lines tighter: one widely cited vendor page, as retrieved for this article, says under about 700 ms feels natural and above about 900 ms callers start talking over the agent. Whichever line you use, the mechanism is the same and the budget is small.

The mechanism

Voice-to-voice latency is a sum. Each stage below runs in series; nothing later can start until the earlier stage finishes.

StageWhat happensWhere the time goesNote
End-of-turn detectionThe speech model decides the caller has finishedFrom a fraction of a second to more than a secondA fixed silence timeout spends its full value on every turn; a streaming end-of-turn model can fire during the last word
Transcript deliveryThe final transcript reaches the language modelTens of millisecondsRarely the problem
Model call 1Decides what to say, or which tool to callRoughly 0.7 s per call in the research team's demo tracesGrows with prompt length and output length
Tool executionYour system answersAnything from tens of milliseconds to secondsEntirely under your control
Model call 2Turns the tool result into speech textRoughly 0.7 s againPresent only on tool turns
Time to first audioText-to-speech emits the first playable sampleRoughly 75 to 313 ms across providers in one reported June 2026 comparisonMeasure playable audio, not the first byte
Playback bufferingThe client holds audio to avoid stutterZero to about 0.8 sTrades smoothness for delay; invisible in server-side numbers

Four consequences follow.

Tool turns are the turns that matter. A plain turn is one model call. A tool turn is two, plus your system's response time. In the research team's demo traces, two model calls alone took roughly 1.4 s before any speech began. Every extra tool the model chains inside one turn adds another call; see fewer tool calls per turn.

End-of-turn detection is the largest single variable. A fixed silence timeout of, say, 700 ms adds 700 ms to every turn before anything else starts, and still cuts callers off when they pause mid-sentence. A streaming end-of-turn model predicts the end from the audio itself and usually fires faster on complete sentences. The fixed silence timeouts anti-pattern covers the trade.

First byte is not first audio. Time to first byte can include container headers with no sound in them. Time to first audio is when a playable sample arrives. The reported source above draws exactly this distinction; ask which one you are being quoted.

Buffering is real latency. On some network paths the agent's speech arrives in clumps at roughly real-time pace. A client that starts playing at once stutters; a client that holds about 0.8 s of audio plays smoothly and adds 0.8 s. The vendor's server never sees this delay, and neither does a test harness that does not play audio.

A worked budget makes the line concrete. A well-built plain turn: end-of-turn 200 ms, one model call 400 ms, first audio 150 ms, buffer 50 ms, total 800 ms. The same build on a tool turn: 200 + 400 + 100 (your lookup) + 400 + 150 = 1,250 ms. That is why tool turns cross the line even when everything is done well, and why the design choices in the next two articles exist.

Evidence

The research team's demo-build measurements are one set of measurements, not a benchmark. Method: scripted callers with pre-recorded audio ran the same flow in a real browser against the same language model and the same tools; voice-to-voice was measured at the client from the end of caller audio to the first agent audio; the median is over the tool-backed turns of a session.

  • Streaming end-of-turn speech models measured roughly 1.5 to 2.0 s voice-to-voice median. A fixed-endpointing configuration with a 700 ms silence timeout measured 2.3 to 3.2 s on the same flow. The difference is close to the timeout plus a slower first turn.
  • Model calls on tool turns measured roughly 0.7 s each. A flow that chained a second tool inside the same turn added about a second; folding the second lookup into the first result removed it.
  • Routing the agent through a distant region added roughly 0.4 to 0.9 s per turn.
  • The first turn of a fresh session was consistently slower than later turns. Warm the system before you time it, and time it again cold.
  • A fixed playback lead of 300 to 450 ms still stuttered dozens of times per call on a jittery link; a content-based prebuffer of about 0.8 s removed the stutter at the cost of that delay.

The reported TTFA range of roughly 75 to 313 ms across providers (June 2026) shows that text-to-speech is usually the smallest term. When a vendor's explanation for slow turns is "the voice", ask for the breakdown.

How to test for it in a demo

  1. Time it yourself. Record the call on your phone. In any audio editor, measure from the end of your last word to the agent's first sound. Do it on ten turns, half of them turns that need a lookup. Write down the median and the worst.
  2. Compare the greeting with a lookup. Ask a question that needs the schedule or an account. If the lookup turn is more than a second slower than the greeting, ask why.
  3. Ask for the breakdown. End-of-turn, model, tool, first audio, per turn. A team that has instrumented the pipeline will have it in a trace. A team that has not will give you one number.
  4. Ask which number it is. First byte or first playable audio. Server-side or at the client.
  5. Call from the right network. Mobile, not the vendor's office wifi. If your callers are in another country, call from there or have someone do it.
  6. Listen for fillers. "Let me check that for you" can honestly cover a tool wait. It becomes a problem when it talks over you or arrives after you have started speaking. Note it separately from latency; see spoken fillers that talk over callers.

A vendor who insists on running the demo from their own recordings has controlled every term of the sum. See vendor-run demos.

Questions to ask vendors

The three questions in the frontmatter cover the sum: tool-turn latency with method, how end-of-turn is detected, and whether the voice number is first byte or first audio. If the answers are a single number, a fixed timeout and first byte, you have learned what the caller will feel.

Questions to ask vendors

  1. 01

    What is your median and 90th-percentile voice-to-voice latency on a turn that calls a tool, and how did you measure it?

    A good answer: Separate numbers for tool turns and plain turns, measured from the end of caller speech to first playable audio on a real phone or browser, with the method stated.

  2. 02

    How does your end-of-turn detection work: a fixed silence timeout, or a model that predicts the turn has ended?

    A good answer: A model-based or adaptive method with its fallback stated; if a fixed timeout, its value and the reasoning behind it.

  3. 03

    When you quote text-to-speech latency, is that time to first byte or time to first playable audio at the client?

    A good answer: Time to first playable audio, measured where the caller hears it, with buffering included.