Skip to content
Voice AgentBible

The verification ladder: from unit tests to a live call on the buyer's network

Green unit tests are not evidence a voice agent works. Only a real browser or phone, golden recordings and a live call on the audience's network catch audio faults.

By · 5 min read

Last verified 30 Sept 2026v1.0Published 30 Sept 2026

Best practiceEvaluationimpact high

Unit tests and scripted harness runs cannot hear audio; a voice agent is verified only after a real browser or phone run, golden recordings, and one live call from a network like the audience's.

Why buyers care

A voice agent can pass every automated test and still be unlistenable. In the research team's demo builds, one hand-over shipped after every unit test passed, every scripted caller clip round-tripped cleanly through the speech model, and end-to-end harness takes passed. The reviewer on the receiving end reported a breaking voice, the caller talking over the agent, and slow tools. All three were real.

The root cause was not a bug in the agent. It was a gap in the evidence. The test harness simulated the browser's timing but never played audio, so playback stutter could not appear in it; its rule for "the agent has finished speaking" was stricter than the browser's, so the caller never overlapped in the harness and did overlap in the browser; and the presenter's network, on another continent from the agent, added jitter the harness never saw.

Buyers care because the same gap exists in most vendors' evidence. A green test suite proves the guards work. It does not prove the call sounds right on your callers' phones.

The mechanism

Verification is a ladder. Each rung sees something the rung below cannot, and each is cheap compared with the rung above. Skipping a rung is where the misses come from.

RungWhat it runsWhat it can seeWhat it cannot see
1. Unit testsGuards, tool registry, state machineLogic errors, guard regressionsAnything about audio or timing
2. Clip round-tripsEach scripted caller line through the speech model actually usedWords the speech model mis-hears, lines that finalise lateTurn-taking, playback
3. Scripted end-to-end harnessThe full pipeline with simulated client timing, one take at a timeTool sequencing, gate behaviour, per-turn latency at the serverPlayback stutter, real overlap, network jitter
4. Real browser or phoneThe actual client page or a real call, instrumentedUnderruns, overlapping audio, caller turns heard versus spoken, what the client waited forThe audience's network
5. Golden recordingsThe verified take, marked and kept; superseded takes deletedExactly what was verified, replayable offlineAnything that changed since
6. Deployed smoke and live listenThe deployed URL or number, one live play, a human listening from a network like the audience'sJitter, cold starts, stale caches, what the room will hearOnly the next network
7. Hand-over noteWhat replays offline, expected latency, what is mockedSets the reviewer's expectationsNothing; this is the receipt

Rules that make the ladder work.

One take at a time. Running six or seven end-to-end takes concurrently produced overlapping turns, split turns and a stall in the research team's builds. The results were not wrong about the agent; they were wrong about everything. Serialize, and check that nothing else on the machine or the shared credentials is running a take.

Instrument the real client. The browser run has to report numbers, not a feeling: audio underruns, overlapping audio chunks, and caller turns heard compared with lines spoken. The passing line is zero underruns, zero overlaps, turns heard at least equal to lines spoken. Keep the full log; a filtered log cost an hour of guessing in one build.

Know what "finished speaking" means in the client. The client should start the next caller line only when the agent has spoken after its last tool event, the audio-complete event has arrived, everything queued has been heard, and a short quiet window has passed. Waiting on any one of those alone produced overlap.

Wait for playback, not arrival. The greeting's audio arrives in about a second and plays for eight to twelve. A rig that starts the caller on arrival talks over the greeting every time.

Mark goldens and delete the rest. A golden is a take that passed the real-client rung. Superseded and partial takes are deleted so the demo's replay picker shows exactly one per flow, and the golden is what replays offline when the room's network is bad.

Listen yourself, from the right place. Watch the rendered clips. Then do one live play on the deployed URL from a network like the audience's, and listen. A stale browser tab running old client code was a real cause of "still choppy" in one build; a build stamp in the status line tells you which code the tab runs.

Evidence

The research team's five demo builds in 2026 supply the pattern. Method: each rung was run in sequence on the same build; misses were recorded when a later rung, or the reviewer, found a failure an earlier rung had passed.

  • A build passed all seventeen unit tests, every clip round-trip and end-to-end harness takes, then failed in the reviewer's hands on voice breaks, overlap and tool latency. The harness could not play audio and its reply-complete rule differed from the browser's.
  • In the real browser, waiting only for the audio-complete event or only for a quiet window produced 409 overlapping audio chunks on four of eight caller lines in one run. The combined rule produced zero.
  • On a jittery link the agent's speech arrived in 152 clumps for 21 sentences at roughly real-time pace. A fixed playback lead of 300 to 450 ms still stuttered 34 times per call; a content-based prebuffer holding about 0.8 s (raised on a dry queue, capped at about 1.6 s) produced zero underruns.
  • Six or seven concurrent harness takes produced overlaps, split turns and a stall that none of the serial takes reproduced.
  • Three clean takes on a slower regional endpoint were rejected by a stale harness assertion after a tool's behaviour changed. The assertion was wrong, not the takes; read the failing assertion before re-recording.
  • The presenter heard a previous day's caller voice after a re-record because the client cached clips for a day. A smoke test on the deployed URL from a fresh browser found it.

How to test for it in a demo

A buyer cannot run the vendor's ladder, but can ask for its artefacts and climb the top rungs personally.

  1. Ask for the artefacts. Test output, a real-browser stats line for the flow you are about to see, the golden recording, and the log of the last live call. A vendor with a ladder has all four within reach.
  2. Call from your own phone on your own network. Not the vendor's office wifi, not their softphone. If your callers are abroad, have someone call from there. Compare what you hear with the vendor's golden of the same flow.
  3. Ask for the numbers from the real-client run. Underruns, overlaps, turns heard. If the answer is "it sounded fine", there is no rung four.
  4. Ask what replays offline. A vendor who cannot fall back to a golden when the room's network misbehaves has not verified what you are hearing.
  5. Ask about concurrency. How many takes at once. The right answer is one.
  6. Ask for the story. The last time a green run shipped a broken call, and what rung was added. Every team that has shipped voice has one.

If the vendor insists on running the demo from their own recording end to end, that is a rung-five artefact presented as a rung-six result. See vendor-run demos and harness-only testing.

Questions to ask vendors

The three frontmatter questions: what the ladder looks like and which rung plays audio, whether golden recordings exist and can be diffed against a live call, and how many takes run at once plus the story of the last miss. The answers tell you whether the vendor's evidence is about the code or about the call.

Questions to ask vendors

  1. 01

    What does your verification ladder look like before a release reaches a customer, and which rung actually plays audio through a browser or phone?

    A good answer: A named sequence ending in a real-browser or real-phone run with instrumented playback stats and a live listen from a representative network, not only unit and harness tests.

  2. 02

    Do you keep golden recordings of verified calls, and can we diff a live call against one?

    A good answer: Yes, with timing; superseded recordings are deleted, and the golden is what the demo replays when the network is bad.

  3. 03

    How many test takes do you run at once, and what happened the last time a green test run shipped a broken call?

    A good answer: One at a time, because concurrent takes corrupt each other's timing; and a specific story with the rung that was added afterwards.