Skip to content
Voice AgentBible

Speech-to-text (STT/ASR)

Speech-to-text in a voice agent: what it does, why accuracy on telephone audio in your languages matters, and what to ask vendors who sub-process it.

By · 1 min read

Last verified 30 Sept 2026v1.0Published 30 Sept 2026

Glossary

The component that turns the caller's speech into text for the language model. Its accuracy on your accents, languages and telephone audio caps everything the agent does afterwards.

Also called: ASR, automatic speech recognition, speech recognition, transcription engine, STT.

What it is

Speech-to-text (STT), also called automatic speech recognition (ASR), converts audio into text as the caller talks. In a voice agent it runs in streaming mode: partial results arrive while the caller is still speaking, and a final result is committed when the system decides the caller has paused or finished. The language model only ever sees this text. A misheard digit, name or drug is a misheard digit all the way through the call.

Telephone audio is narrow-band and often compressed, so results measured on clean studio recordings do not transfer. Accents, code-switching and background noise widen the gap further.

Why it matters when buying

Many voice-agent vendors license speech recognition from a third party rather than build it. That is normal, but it has consequences: data-processing terms, the ability to add your vocabulary, the languages actually supported at production quality, and who is accountable when accuracy drops. Speech-recognition systems can also produce text when there is no speech at all (see hallucination), which turns silence into false intents.

What to ask

Ask which speech engine is in the audio path and whether it is disclosed as a sub-processor. Ask for word error rate on telephone audio in your languages, with the test set described. Ask how custom vocabulary (product names, staff names, local place names) is added. Ask what the agent does with an empty or noise-only turn.

← All terms