Harness-only testing: your test rig cannot hear audio
A scripted harness passes every take while the real call stutters and overlaps, because it never plays audio. Why a real-client run with counters is mandatory.
By Voice Agent Bible Research · 5 min read
Last verified 30 Sept 2026v1.0Published 30 Sept 2026
A text-and-timing harness cannot observe playback stutter, overlap or network jitter; the only evidence of a working voice agent is a real client on a realistic network with counters.
Symptoms a buyer notices
- Every automated test passes, and the first human listener reports broken voice, the caller talking over the agent, or long tool waits.
- The vendor's evidence of quality is a transcript and a latency table, never a recording made from a real phone or browser.
- Problems appear only from the customer's location or network and cannot be reproduced from the vendor's office.
- A fix is declared from a passing harness run without anyone having listened to the result.
How it shows up
The build passes. Every unit test is green, every recorded audio clip round-trips through speech recognition cleanly, and the end-to-end scripted harness prints a pass on every take. The demo is handed over. The first person to actually listen reports three things: the agent's voice breaks up, the caller talks over the agent, and the tool turns feel slow. None of that appeared in any test, because none of the tests could hear.
The pattern repeats wherever a voice product is verified by a text pipeline. A test rig sends caller audio or caller text, waits for the agent's reply as a transcript or an event, checks the words and the timing, and moves to the next line. It never decodes the agent's audio into a playback buffer, never schedules that buffer against a wall clock, and never runs over the network the audience will use. The categories of failure that live in those three places are simply not observable from inside the rig.
Why buyers care
Callers judge a voice agent by how it sounds in the first ten seconds. Stutter, clipped words, a greeting that is interrupted by the agent's own next sentence, or the caller and agent speaking at once are the failures people notice before they notice whether the answer was right. A vendor whose entire quality evidence is transcripts and latency tables has proven that the words were right. That is a different claim from "the call was listenable".
Network is part of the product. A vendor tests from an office with a good connection close to their servers. Your callers dial from a mobile handset in another country, or your presenter runs the demo over hotel wireless. Audio that arrives in a smooth stream on the vendor's network arrives in clumps on yours, and clumps are what a playback buffer turns into stutter. A harness on the vendor's network cannot see this; a real client on your network cannot miss it.
The mechanism
A voice agent's reply reaches the listener as a stream of small audio chunks. Three things must go right for it to sound like speech: the chunks must arrive fast enough to stay ahead of playback, the client must decide when to start playing so that it does not run dry, and the client must know when the reply is finished so that the next thing (the caller's line, a new prompt) does not land on top of it. A harness that treats the reply as "the transcript arrived" or "the audio-done event arrived" has abstracted away all three.
Arrival is bursty. On some links the agent's speech arrives at roughly real-time pace in clumps. In one measured run, 21 sentences arrived as 152 separate arrival clumps. A player that starts on the first chunk with a fixed small lead runs dry every time a clump is late, and each dry moment is an audible stutter.
Start rules matter more than lead time. A fixed lead of 300 to 450 milliseconds still stuttered 34 times per call on such a link. What worked was content-based: hold the reply until about 0.8 seconds of audio is queued (or the reply is complete, or 1.2 seconds have passed), then play contiguously, appending later chunks gaplessly, and if the queue still runs dry mid-reply, raise the prebuffer in 0.3-second steps up to a cap. None of this logic exists in a text harness, so none of it can be tested there.
Completion rules decide overlap. The harness's rule for "the agent has finished" was stricter than the browser's, so the harness never sent the next caller line early and never produced an overlap. In the real client, waiting only for the audio-done event, or only for a quiet window, released the next caller line over the tool-backed second sentence. Measured: 409 overlapping audio chunks on four of eight caller lines. The rule that fixed it needs four conditions at once: spoken text after the last tool event, that speech's audio-done, everything queued has been heard, and 1.2 seconds of no new audio.
Greetings arrive fast and play slowly. A greeting's audio arrived in about one second and played for 8 to 12 seconds. A rig that starts the first caller line on arrival makes the caller interrupt the agent's introduction every time, and a presenter is the first to notice.
Stale clients. A fix to the player is invisible if the tab or app under test is still running the old build. A build stamp in the visible status line was what made "still choppy" reports diagnosable.
The general point: a harness simulates the client's timing model. The client's timing model is the thing under test.
Evidence
From the research team's demo builds, method on the methodology page:
- A build shipped after 17 of 17 unit tests, clean clip round-trips and passing end-to-end harness takes, and the reviewer reported broken voice, caller-over-agent overlap and slow tool turns. Root cause: the harness never played audio and its wait-for-quiet rule was stricter than the browser's. The presenter's intercontinental network added jitter the harness never saw.
- Real-client instrumentation exposed counters for underruns, overlapping chunks, overlapping lines, voice-to-voice per turn and completion. The required gate result is zero underruns, zero overlapping chunks and at least as many caller turns as scripted lines. Runs that fail any of the three are not shipped.
- Fixed-lead playback: 34 stutters per call on a bursty link. Content-based prebuffer: zero underruns on the same link in the gate runs.
- Completion rule: 409 overlapping chunks on four of eight lines with an audio-done-only rule; zero overlapping chunks with the four-condition rule.
- Six or seven concurrent harness takes on one account produced overlaps, split turns and a stall that never appeared in serial runs; another session's gate on the same account did the same. Concurrency in the test environment is itself a variable that must be controlled.
- A grep-filtered run log cost an hour of guessing; the full timeline was the evidence. The gate saves its complete output.
The final rungs of the research team's verification ladder are therefore a real-client run with counters and then one live play on the deployed address, on a network like the presenter's, with a human listening. No handover happens without both.
How to test for it in a demo
- Ask for the counters. Underruns, overlapping chunks, user turns, voice-to-voice per turn, from the vendor's last release gate, with the client and network named. A vendor who has them answers in a minute. A vendor who offers a transcript instead has a harness.
- Call from where your callers are. Use a mobile handset on a cellular network in the region your callers dial from. Watch the live transcript while you listen. Count stutters, clipped words and moments where you and the agent spoke at once. Ten turns is enough.
- Interrupt the greeting. Start talking two seconds into the agent's introduction. Pass: it stops cleanly and takes your turn. Fail: it talks over you or restarts.
- Listen to a recording from the listener's end. A server-side mix hides playback problems because the server never had them. Ask for a client-side capture of a recent production call.
- Check the build. Ask how the vendor knows the client under test is running the current build. A visible build stamp is the good answer.
- Ask how a remote complaint is reproduced. The good answer is running the real-client gate from that region or through an emulated network profile and then listening. Re-running the text harness is the anti-pattern.
Questions to ask vendors
- How do you test that the agent's speech plays without stutter or overlap, on what client, and on what network? A good answer is a real client driven automatically, with counters that must read zero, from a network like the callers', before every release.
- Can you show me the counters from your last release gate, and the full timeline rather than a summary? A good answer is a saved run with underruns, overlapping chunks, user turns and per-turn latency.
- When a customer in another region says the voice breaks, how do you reproduce it? A good answer involves that region's network and a human listening to the recording.
- Who listened to the last release before it shipped, from where, and what did they hear? A vendor who cannot name a person and a network has not listened.
Questions to ask vendors
- 01
How do you test that the agent's speech plays without stutter or overlap, on what client, and on what network?
A good answer: A real client (browser or phone) driven automatically, with playback underrun and overlap counters that must read zero, run from a network like the callers' before every release.
- 02
Can you show me the counters from your last release gate, and the full timeline, not a summary?
A good answer: Yes, with underruns, overlapping audio chunks, user turns and voice-to-voice per turn, from a saved run.
- 03
When a customer in another region says the voice breaks, how do you reproduce it?
A good answer: By running the real-client gate from that region or through an emulated network profile, then listening to the recording, not by re-running the text harness.
Related
- Best practice
- Best practice
- Best practice
- Best practice