Words from nothing: speech recognition fabricating text on silence and noise
Speech recognition invents words from silence, comfort noise and dither. A probe any buyer can run, and the timestamp signature that catches most fabrications free.
By Voice Agent Bible Research · 6 min read
Last verified 30 Sept 2026v1.0Published 30 Sept 2026
Every speech-recognition system tested fabricates words on speech-free audio at some rate; a word whose timestamps fall outside the audio is fabricated by construction.
Symptoms a buyer notices
- Transcripts contain a greeting, a yes, a digit or a product name on a channel where nobody was speaking.
- The agent acts on a confirmation the caller never gave, usually right after a pause, a hold tone or line noise.
- A term from your custom vocabulary list appears in a transcript of a moment when nothing was said.
- A call in one language produces a short fluent phrase in another.
How it shows up
A transcript of a call contains a word nobody said. Often it is a small word: a greeting, a yes, a digit, a filler. Sometimes it is a specialised term, a drug name or a product name, sitting in the middle of what the recording shows was a hold tone. Sometimes it is a short phrase in a language the caller was not speaking. The word carries a confidence score, it appears in the final transcript, and downstream systems treat it as real.
In a voice agent the symptom is behavioural rather than textual. The caller pauses to find a card, the line carries comfort noise for two seconds, and the agent says "great, I'll go ahead and book that". Nobody said great. Or an agent capturing a reference number records an extra digit that arrived during a breath. Or a call that should have been silent on one channel (a hold queue, a ringing line) produces an analytics record of a conversation that never happened.
Buyers usually meet this as a mystery ticket: "it transcribed a word we never said, and it is one of our custom vocabulary terms."
Why buyers care
Speech recognition is the first component in the chain, and everything after it trusts its output. A fabricated affirmative is indistinguishable, to the language model, from a real one. A fabricated digit in an account number, a one-time code or a date of birth is a wrong write into a system of record. A fabricated phrase in a monitored call is a compliance record of something that did not occur.
The problem is invisible in ordinary accuracy testing. Word error rate is measured on audio that contains speech. Speech-free audio is exactly the case that accuracy benchmarks exclude, and exactly the case that every real phone call contains: the seconds before the caller speaks, the pauses, the hold, the wrap-up silence after goodbye.
The mechanism
Speech-recognition models decode audio in windows. When the audio is shorter than the window, or ends partway through the last window, the remainder is padding. The model still has to emit something for that region, and its training has taught it what usually follows quiet audio: a greeting at the start, an acknowledgement at the end, whatever tokens were common in the quiet stretches of its training data. On speech-free audio the model has nothing to anchor to and reaches for those habits.
Four properties follow, and all four were reproduced in the research team's study.
Short clips are worse. The shorter the audio, the larger the share of the final window that is padding rather than sound. Clips of 1.3 seconds and under fabricated at roughly three times the rate of 30-second clips.
Each model has a small vocabulary for nothing. Fabrications are not random. A given model reaches for the same handful of tokens across unrelated recordings: a greeting, an affirmative, a filler, and for domain-tuned models, domain terms. A general-purpose model that says "hi" to noise and a medical model that says a drug name to the same noise are showing the same mechanism with different training data.
Vocabulary steering can choose the invented word. Custom vocabulary features raise the likelihood of listed terms. On speech-free audio that boost still applies. In the study, a recording that reliably fabricated one domain term fabricated a different, phonetically similar listed term instead, six times out of six, when that term was supplied as a custom vocabulary entry. Phonetically distant terms never captured (zero of eight control cells). Every custom vocabulary list is therefore a candidate fabrication vocabulary for the quiet parts of the call.
Language detection on a speechless channel is decided by your configuration. When a channel carries no speech and the system is asked to pick from a list of candidate languages, it still picks. In the study it committed to a language on every speechless channel tested, at a reported confidence of zero. Which language it commits to is decided by the candidate list, not by the audio. If that language's model then fabricates, the result is a fluent phrase in a language nobody on the call was speaking.
The timestamp signature. Because the invented tokens are decoded from padding, most of them carry word timestamps at or beyond the end of the audio. A word that starts three seconds after a 30-second file ends, or a word spanning three seconds on a 0.3-second clip, cannot have come from sound. This is the cheapest detector available: compare each word's end time with the audio duration and drop or flag anything past it. It held across all three architectures tested, including an open-weight family known for this behaviour.
Evidence
The research team's probe study, method on the methodology page, sent roughly 6,500 requests of speech-free audio to production speech-recognition services covering three model architectures. The audio was synthesised from logged seeds so that every clip is byte-reproducible and contains no speech by construction, which makes any returned word provably fabricated.
- On 500 fresh 30-second realisations of low-level pink noise per model, two commercial models fabricated a word on 2 to 3 percent of recordings. The same clips through an open-weight model returned nothing at 30 seconds but fabricated a single filler word on essentially every clip under 1.3 seconds.
- Pooled across the pink-noise family, clips of 1.3 seconds and under fabricated on 6.2 percent of calls versus about 2.4 percent at 30 seconds.
- Fabricated words on noise scored word confidence of 0.54 and below, while real speech in the same pipeline sat above 0.9. A confidence floor near 0.6 therefore catches the noise class. It does not catch the wrong-language class, where fabricated words reached 0.98.
- Fabrication survived the real telephony chain. Re-encoded through 8 kHz companded telephony codecs, a mid-single-digit share of calls still fabricated; two narrowband compressed codecs suppressed some but not all of it.
- Exact digital zeros never fabricated in any condition tested. Telephony comfort noise at very low level did. The difference matters: a channel that is muted to true zeros is safe in a way that a channel filled by the carrier's comfort-noise generator is not.
- Six classes of audio engineered to sound speech-like (syllable-rate gating, formant-band noise, speech-rate modulation) produced nothing. Plain unstructured noise is the trigger, not speech-likeness.
- Fabrication is stochastic per request. Identical bytes returned different outputs across repeats, and per-recording propensity ranged from roughly 15 percent to 100 percent across six-repeat re-tests. A single clean probe run proves nothing; six repeats is the minimum.
Two apparent findings in the study were killed by checking the instrument rather than the reading: a dramatic neighbour-channel effect turned out to be the harness altering the audio, and a wrong-language detection rate on real speech turned out to be truncated test clips. Both are worth mentioning because a buyer running this probe will meet the same traps.
How to test for it in a demo
This is a probe any buyer can run against any pipeline, including through a live phone call.
- Build speech-free audio. Sixty seconds in four segments: true digital silence, telephony comfort noise (or a recording of a muted handset on a real line), a ringing or hold tone, and low-level broadband noise. Also cut ten clips of 0.3 to 1.3 seconds from the noise segment.
- Send it through the pipeline the agent actually uses, not a separate transcription endpoint, with the vendor's production settings including any custom vocabulary. Repeat each clip six times.
- Count every word returned. Zero is the only acceptable count. Record each word's confidence and its start and end times.
- Check timestamps against duration. Any word ending after the audio ends is fabricated by construction. Ask whether the vendor filters these before the transcript reaches the agent.
- Swap the vocabulary list. Replace the custom terms with phonetic neighbours and repeat. If the invented word changes, vocabulary boosting is steering the fabrication.
- Test the agent, not just the transcript. Ask the agent a yes-or-no question, then send two seconds of comfort noise. Pass: it waits or asks again. Fail: it proceeds.
If the vendor cannot let you inject audio, do it by phone: call the agent, ask for a confirmation question, mute your handset for five seconds, and see what it does. Repeat ten times.
Questions to ask vendors
- What does your speech-recognition pipeline return on a channel that carries only comfort noise for a minute, and how did you measure it? A good answer is a rate per minute with a method, plus a description of the filters between recogniser and agent.
- Do you filter or flag words whose timestamps run past the end of the audio, and words below a confidence floor? A good answer gives the thresholds and shows the rejected words in a log.
- How does a short, quiet backchannel become a confirmation in your agent? A good answer describes a confirmation guard that needs a plain answer to a specific question and treats single-word turns on noisy audio as suspect.
- If we supply a custom vocabulary list, what stops those terms appearing in silent stretches? A good answer acknowledges the mechanism and describes the mitigation; a blank look is a finding.
This is a property of the category, not of one product. Every architecture tested did it at some rate, and the rates moved with clip length, noise character and configuration rather than with the vendor's logo.
Questions to ask vendors
- 01
What does your speech-recognition pipeline return on a channel that carries only comfort noise for a minute, and how did you measure it?
A good answer: A measured fabrication rate per minute of speech-free audio, the method, and the filters applied before a transcript reaches the agent.
- 02
Do you filter or flag words whose timestamps run past the end of the audio, and words below a confidence floor?
A good answer: Yes to both, with the thresholds stated and the rejected words logged so you can audit them.
- 03
How does a short, quiet backchannel such as 'yes' or 'mm-hm' become a confirmation in your agent?
A good answer: It does not on its own; a confirmation needs a plain answer to a specific question, and single-word turns on noisy audio are treated as suspect.
Related
- Best practice
- Best practice
- Best practice