Skip to content
Voice AgentBible

Time to first audio (TTFA)

Time to first audio (TTFA) defined for voice-agent buyers: what it measures, how it differs from voice-to-voice latency, and how to get a number you can trust.

By · 1 min read

Last verified 30 Sept 2026v1.0Published 30 Sept 2026

Glossary

How long after the agent has something to say the caller hears the first sound. Usually quoted for the text-to-speech step alone; buyers should also ask for the end-to-end version measured from the end of caller speech.

Also called: TTFA, time to first byte of audio, first-audio latency, TTS latency.

What it is

Time to first audio is the delay between a text-to-speech request starting and the first audio sample being available to play. Speech vendors quote it as a component metric, often in the low hundreds of milliseconds. It measures one link in the chain: the synthesiser's own responsiveness.

The caller experiences something longer. Before synthesis can begin, the system must decide the caller has finished speaking, transcribe the last words, wait for the language model's first tokens and, on many turns, complete a tool call. The end-to-end version of this delay is voice-to-voice latency, and it is the number that determines whether the conversation feels natural.

Why it matters when buying

Vendors sometimes quote TTFA where a buyer hears "response time". A synthesiser that starts in 150 milliseconds inside a pipeline that takes 1.5 seconds to reach it is not a fast agent. Ask which quantity is being quoted, from which point to which point, and measured where (in a browser on the vendor's network, or over a real phone call).

What to ask

Ask for the definition in writing: start event, end event, measurement point. Ask for median and 90th percentile, not a single "typical" figure. Ask for the numbers on a turn that calls your system, because greetings are fast everywhere. The methodology page describes how this site measures it.

← All terms