How to run a voice agent bake-off: a 90-minute protocol buyers control
A step-by-step protocol for comparing AI voice agents on your own audio, your own traps and your own sandbox, with a stopwatch latency test and a scoring sheet.
By Voice Agent Bible Research · 8 min read
Last verified 30 Sept 2026v1.0Published 30 Sept 2026
A bake-off is a controlled comparison of two or more voice agents on calls you design, from a phone you hold, into a sandbox you own. Enterprise buyers run it before a pilot. Small businesses can run a lighter version in an afternoon. Either way the aim is the same: replace the vendor's demo with your evidence. This guide takes about 90 minutes per vendor once the preparation is done. The scoring sheet, the adversarial personas and the industry scripts are the companion assets.
1. Why vendor demos mislead
A vendor demo is built to succeed. That is not dishonest, it is the job. But understand what you are watching.
Scripted happy paths. The caller in a demo says exactly what the agent expects, in the order the agent expects it. Real callers give three facts in one breath, change their mind mid-sentence and answer a question that was never asked.
Cached audio. Demo callers are often pre-recorded and pre-tested. Every line has been transcribed in advance and reworded until it comes out clean. Some demos replay a whole recorded session with its original timing, so nothing touches the network at all. A recorded demo is a useful safety net for a presenter and a useless measure of a product.
A cherry-picked network. The demo runs from a warm machine, in the vendor's chosen region, on a wired connection. Your callers arrive on a mobile network through a carrier, then through your telephony provider. Latency and audio quality both change.
The vendor at the keyboard. The presenter knows which words trigger which tool and which ones to avoid. They know the one turn that sometimes stalls and they talk over it. They know the recorded caller's voice may be a day stale and the fix is a page refresh.
None of this tells you how the agent behaves with your callers. So you change all four conditions.
2. Before the demo
Preparation is most of the work. Budget a half day.
Five scenarios. Pull them from your real call mix, not from the vendor's list. For a clinic: a new booking, a reschedule, an insurance question, an emergency, a wrong number. For a collections team: right-party contact, wrong number, dispute, promise to pay, callback request. Each scenario has a defined "done" state that you can observe in your own system.
Three personas plus one adversarial. Three callers who behave normally but differently: one brisk, one slow and chatty, one with an accent the vendor did not choose. One adversarial caller whose job is to break the agent. The adversarial personas page has eight ready-made cards.
A KPI sheet. Pass criteria written down before anyone calls. Voice-to-voice latency target, containment definition, transfer rules, the disclosures that must be spoken aloud. The scoring sheet is the template.
Your own recorded caller audio. This is the single biggest change you can make to a demo. Record a colleague reading each persona's lines and hand the files to every vendor, or play them into the phone yourself. Four rules make the recordings fair and reproducible:
- One sentence per line. A full stop and a pause in the middle of a line splits it into two turns on most turn-taking systems. Join clauses with commas or "and".
- End every line on a word, never on a digit. A line that ends "...seven, seven, zero, three, one" is often finalised seconds late, because the recogniser waits to see whether more digits follow. End it "...zero, three, one, for the account in my name."
- The voice matches the script. If the line says "this is Rosalind", the recording is a female voice. If the caller is meant to be from a region, the accent is from that region.
- Digital silence between lines. Do not record room noise between sentences. Some endpoint detectors hold a turn open on a synthetic noise floor and finalise late. Silence is the neutral condition; you will add noise deliberately in the trap ladder.
Run every clip through a speech recogniser of your own choosing before the bake-off and fix garbles by rewording, not by hoping. If a line does not transcribe cleanly for you, it is a poor test of the vendor.
Finally, insist on a sandbox of your own system of record, populated with fictional records that you supply. More on that in step 7.
3. The latency test with a phone and a stopwatch
You do not need tooling to measure latency. You need a phone, a stopwatch and patience.
Define two numbers:
- Time to first audio (TTFA). From the moment you stop speaking to the first sound the agent makes, including a filler such as "one moment".
- Voice to voice. From the moment you stop speaking to the first meaningful word of the answer.
Fillers make TTFA look good and leave voice to voice unchanged. Measure both, and score on voice to voice.
Measure tool-backed turns separately from conversational turns. A greeting and a "yes, I can help with that" are fast on every platform. A turn that checks availability, looks up an account or files a record involves at least one round trip to a tool and often a second reasoning pass, so it is the turn that decides whether the agent feels usable. Ask the vendor which turns call a tool, then time those.
A widely repeated rule of thumb, taken here from a vendor's buyer guide and treated as a vendor claim rather than a measurement, is that under about 800 ms per turn feels conversational and above about 1.2 s starts to feel like an IVR. Use it as a frame, not as a verdict. Your callers' tolerance depends on the task: a two-second pause before a booking confirmation is fine; a two-second pause before every "yes" is not.
Practical method: one person calls and speaks, another holds the stopwatch and starts it on the last syllable. Take at least ten tool-backed turns per vendor and record the median and the worst case. Do this from a mobile phone in the room, and again from a landline or a softphone if that is how your callers arrive. Note the first turn of a fresh session separately; cold starts are common and worth knowing about, but they are not the steady state.
4. The trap ladder
Run the same traps against every vendor, in the same order, with the same recorded audio where possible. Each trap has a pass condition written before the call. Nine traps cover most production failures:
- Barge-in. Interrupt the agent halfway through an offer with a new constraint. Pass: it stops within a word or two and takes the new constraint. Fail: it finishes the sentence, or it stops and then repeats the old offer.
- Eight seconds of silence. Say nothing after the agent asks a question. Pass: one gentle prompt, then a hold or message option. Fail: it hangs up, repeats the whole question, or fills the silence by talking to itself.
- Background noise. Play a recording of a café or traffic under the caller line. Pass: it still gets the intent and asks to confirm anything uncertain. Fail: it treats noise as speech and interrupts, or it endpoints late on every turn.
- Accent. A caller with an accent the vendor did not choose. Pass: intent captured, and any uncertain name or number confirmed. Fail: repeated "sorry, I didn't catch that" loops.
- Number and date read-back. Give a phone number quickly, a date of birth with a mumbled month, and "next Thursday" on a Wednesday. Pass: digit-by-digit read-back and a named date before anything is written. Fail: it writes what it heard.
- Out-of-scope request. Ask something the agent should not do (a clinic agent asked for a diagnosis, a collections agent asked for someone else's balance). Pass: a clear boundary and a route to the right place. Fail: it improvises an answer.
- Angry caller. Raise your voice, use short sentences, repeat a grievance. Pass: it acknowledges, does not argue, and moves to a resolution or a human. Fail: it apologises in a loop or lectures.
- "Let me speak to a human." Pass: an immediate route, with the context passed on, and a stated wait or callback time. Fail: it tries to keep you, or transfers cold.
- Repeat the same request twice within ten seconds. "Book that" then "yes, book that" before the first write finishes. Pass: one record, one confirmation. Fail: two records, or a second confirmation question for the same action.
Add industry traps from the relevant script (emergency language in healthcare, a wrong-number path in collections, a "wait, never mind" at a confirmation in insurance).
5. Scoring
Score before you discuss. Two people score each call independently and reconcile afterwards. Four groups:
- Trap results. Pass or fail per trap, nothing in between. A partial pass is a fail with a note.
- Functional beats. Did the job get done: the record created, the slot booked, the promise captured, the transfer completed. Each beat is observable in your sandbox, not in the vendor's screen.
- Recovery. What happened after each failure. An agent that mishears a number and corrects on read-back scores higher than one that never mishears in a demo, because you saw it recover.
- Disclosures said aloud. The AI disclosure, the recording notice, and any industry-specific statement (a debt-collection disclosure, a consent check on an outbound call). Score them only if you heard them. A disclosure in the prompt that the agent skipped is a fail.
Weights and pass criteria are on the scoring sheet. Keep hard stops separate from scores: a missed emergency instruction or a disclosed balance before identity was verified is a stop, whatever the total.
6. Ask for the logs
After the calls, and before the pricing conversation, ask for:
- Per-call traces for the calls you just made, with timestamps for every turn.
- A latency breakdown per turn, split into speech recognition end-of-turn, reasoning, tool calls and speech synthesis first byte. Compare it with your stopwatch numbers; they should agree within a few hundred milliseconds.
- Tool outcomes. Every function the agent called, its arguments, whether it was allowed, and what it returned. This is where you check that the confirmation before an irreversible action was enforced by code and evidenced in the transcript, not merely suggested to the model.
A vendor who cannot show you this for a demo made ten minutes ago will not be able to show it to you when a caller complains in production.
7. Integration proof
The agent must write into a sandbox of your system, not into the vendor's dashboard. A booking you can see in your own schedule during the call, a promise-to-pay note in your own collections system, a lead in your own CRM. Supply the fictional records yourself so nothing about the test is prepared by the vendor.
Test the warm transfer the same way. When the agent hands the call to a person, that person should receive the caller's name, the intent, what has already been confirmed and what remains. Have a colleague take the transfer and score what arrived. A transfer that starts with "how can I help you?" has lost everything the agent learned.
Ask what happens when the sandbox is slow or unreachable. A good agent tells the caller something honest and offers a callback; a poor one waits in silence.
8. Pricing proof
Advertised per-minute prices are rarely the whole number. Ask for an all-in figure for your volume, broken into line items: platform or agent minutes, speech recognition, speech synthesis, the language model, telephony in and out, and any per-call or per-seat fees.
Then ask how the meter runs on the calls that are not clean conversations:
- Transfers. Does billing stop when the call leaves the agent, or run for the whole call?
- Abandoned calls. Is there a minimum charge for a call that connects and hangs up in a second?
- Silent calls. Sockets that connect and send no speech are a real share of production traffic. Are they billed as minutes?
Reconcile the quote against the calls you made in the bake-off. If the vendor can produce a per-call cost for those, they can produce it for production.
9. Red flags
Any one of these is a reason to pause; two are a reason to stop.
- The vendor insists on their own caller audio or on running the demo from their machine.
- No adversarial persona allowed. "That is not a realistic caller" is not an answer; it is a description of production.
- No sandbox. Every write lands in a dashboard you will not use.
- "Latency depends on your network" with no numbers attached. Everything depends on the network. Ask for the median and worst case on tool-backed turns in their own environment.
- Prompt-only compliance. A disclosure, a consent check or a confirmation step that exists as an instruction to the model and nowhere else. Ask to see where the code refuses the action.
10. Decide with the sheet
Total the weighted scores. List hard-stop failures separately. If two vendors are close, the one with better recovery and clearer logs is usually the safer choice, because production is mostly recovery. Write down the one result that would change your decision, and re-test that one thing on a second day before signing. A bake-off that changes nothing about the vendor's demo has not been run.
Related
- Asset
- Asset