Vendor-run demos: scripted happy paths, cached clips and no adversarial persona
A vendor-run demo is a rehearsal. Why cached audio, scripted loops and friendly personas hide the failures you meet in production, and how to take the demo back.
By Voice Agent Bible Research · 5 min read
Last verified 30 Sept 2026v1.0Published 30 Sept 2026
A demo the vendor drives is a rehearsed recording of the agent at its best; only a demo the buyer drives, with an adversarial persona, measures the agent.
Symptoms a buyer notices
- The vendor plays the caller, uses their own device and network, and follows a script they wrote.
- The caller's lines sound identical each time the demo runs, or the demo is offered as 'a recording of a real call'.
- Every request is answered on the first try; nobody interrupts, changes their mind, mishears or gives a number wrong.
- When you ask to try it yourself, the answer is a follow-up meeting or a sandbox that arrives after the decision.
How it shows up
The vendor shares a screen. A presenter clicks a button, a caller's voice plays, and the agent answers. Every turn lands. The caller never interrupts, never hesitates, never gives a wrong digit and never asks something the agent cannot do. The presenter narrates the tool calls as they happen. At the end, a booking appears in a dashboard. You are impressed, and you have learned almost nothing about how the agent will behave with your callers.
Three details, if you notice them, tell you what kind of demo you are watching. The caller's lines sound the same every time, down to the breath before the first word: they are recordings. The presenter picks from a list of workflows rather than talking: it is a scripted state machine. And when you ask to try it yourself, the answer is a sandbox later, or a follow-up meeting. The demo is a performance, and the vendor is the only performer.
Why buyers care
A voice agent's production behaviour is decided by the cases a rehearsed script avoids. Real callers interrupt, change their minds, correct themselves, mumble numbers, have children in the background, ask two things in one breath and answer questions that were not asked. Those cases are where turn-taking, confirmation logic, stall recovery and barge-in either work or do not. A happy-path script contains none of them by construction.
Recorded or cached caller audio removes even more. If the caller's side is a recording, then the timing of the conversation is partly recorded too, and what you experience as smooth turn-taking may be the replay of a good take rather than the agent responding live. Latency, the thing most buyers say they care about most, is not demonstrated at all by a replay.
Finally, a demo you cannot drive cannot be compared. Two vendors running two different scripts on two different networks with two different presenters produce two performances. A shortlist decided that way is decided by presentation skill. The bake-off guide at /demo-guide/how-to-run-a-voice-agent-bake-off/ exists so that every vendor runs the same script, driven by you.
The mechanism
Demo builds are engineered for reliability in the room, and the same engineering that makes them reliable makes them unrepresentative. The research team has built and handed over several such demos; the traps below were all found in them, which is how we know what to look for.
Recorded callers. Demo builds commonly bake the caller's lines to audio in advance so that the demo does not depend on a presenter's voice, microphone or accent. Those clips are served to the browser like any other asset. In one build they were served with a day-long cache lifetime, so after the clips were re-recorded the presenter heard yesterday's caller for a day. The point for a buyer is not the caching bug; it is that the caller's voice was a file, and a file does not test the agent.
Offline replays. Demo builds also record "golden" takes and can replay them offline with their recorded timing, so that the demo survives a bad hotel network. That is a sensible fallback for a keynote and a misleading artefact in a procurement. A replay demonstrates that the agent once did this; it does not demonstrate that it will do it now, on your line.
Scripted state machines. The caller side of a demo is often a loop that plays line one, waits for the agent to finish, plays line two, and so on. In one build, making the workflow picker clickable during a run without binding each loop to a run token let the previous workflow's loop keep playing its lines into the next workflow's session; a line from one scenario appeared, spoken by the caller, inside a different scenario, with two caller bubbles on screen at once. The bug was fixed. The lesson for buyers is that the "caller" in such a demo is a program with its own state and its own bugs, and its lines were written by the vendor to succeed.
Warmed machines and friendly networks. The first turn of a fresh session is slower; presenters warm the system with a practice run before the meeting. Presenters also run from a good network close to the agent's servers. Your callers do neither.
No adversarial persona. Every scripted caller is polite, patient and correct. The bake-off guide requires three normal personas and one adversarial one for exactly this reason: the adversarial persona is where the product is measured.
Evidence
From the research team's demo builds, method on the methodology page. These are patterns observed while building demos, presented so that buyers can recognise them; none is specific to one vendor.
- Live tool-backed turns measured 1.5 to 2.0 seconds median voice-to-voice on the smoother configuration and 2.3 to 3.2 seconds on the slower one. An offline replay of a golden take has no live latency at all; it reproduces whichever take was kept.
- Greeting audio arrived in about one second and played for 8 to 12 seconds; scripted loops wait for playback to finish before the first caller line so that the recorded caller never interrupts the introduction. A real caller interrupts it about half the time.
- Cached caller clips (a day-long cache lifetime) meant re-recorded lines were not heard until the next day. The clips were re-served without caching afterwards; the underlying fact, that the caller was a file, did not change.
- A scripted loop without a run token played one scenario's caller line into another scenario's session, visible as two caller bubbles at once. Fixed by binding every loop to a run token and checking it after every wait.
- Confirmation guards, stall watchdogs and barge-in gating were each found and fixed only in runs where the caller did something the script did not: a confirmation arriving before the question, a caller line landing on the agent's second sentence, the agent's own voice triggering speech-start through the speakers. Scripted happy-path takes never exercised any of them.
How to test for it in a demo
The remedy is to take the demo back. The full method is in the bake-off guide; the short version:
- You hold the phone. Call the agent from your own handset, on your own network, ideally from the region your callers are in. No vendor device, no shared screen for the audio, no presenter narration.
- You bring the script. Use the site's demo script for your use case or write your own. Include the traps: an interruption mid-offer, a mid-sentence pause, a wrong digit then a correction, two questions in one breath, eight seconds of silence, an emergency or an out-of-scope request. Score each trap pass or fail.
- You bring the adversarial persona. Impatient, hard of hearing, in a noisy room, or trying to get the agent to do something it should refuse. Run this persona last and score it separately.
- Switch off the props. Before you start, ask the vendor to disable cached caller clips, offline replays, presenter mode and any pre-recorded fallback, and to confirm that every turn you hear is generated live.
- Run it twice. Same script, different answers to the agent's follow-ups. Identical agent timing or identical caller audio across the two runs is a replay signature.
- Record it. The recording goes into the scoring sheet alongside the other vendors' recordings of the same script. That is the comparison.
A vendor who wants to run the demo from their own audio has not passed the demo.
Questions to ask vendors
- Can I call the agent now, from my phone, with my own script, and can we record it for the scoring sheet? The only good answer is yes, immediately.
- Is any part of this demo pre-recorded, cached or replayed, and how would I know? A good answer states plainly what is live, labels what is not, and switches the props off for the evaluation.
- Show me the last call where the agent failed. What went wrong and what did you change? A good answer is a specific recent recording, an honest account and a test that now catches the failure. A vendor with no failures to show has not looked.
- Which of your existing customers can I call, unannounced, to hear the agent in production? A production line answered by the agent is the demo that cannot be rehearsed.
Questions to ask vendors
- 01
Can I call the agent now, from my phone, with my own script, and can we record it for the scoring sheet?
A good answer: Yes, immediately, with no vendor-side script or device involved, and the recording goes to you.
- 02
Is any part of this demo pre-recorded, cached or replayed, and how would I know?
A good answer: A clear statement of what is live and what is not; recorded caller audio or offline replays are labelled and switched off for the evaluation.
- 03
Show me the last call where the agent failed. What went wrong and what did you change?
A good answer: A specific recent recording, an honest description of the failure and the fix, and a test that now catches it.
Related
- Best practice
- Best practice
- Best practice
- Best practice
- Use case