Agent-to-agent evaluation: synthetic callers, personas, a rubric judge and honest latency
Evaluate a voice agent with a synthetic caller on the production path, personas, a rubric-scored judge, a per-turn latency breakdown and human spot checks.
By Voice Agent Bible Research · 5 min read
Last verified 30 Sept 2026v1.0Published 30 Sept 2026
Evaluate with a synthetic caller on the production path, personas including hostile callers, a rubric judge spot-checked by people, and latency measured at the client.
Why buyers care
Every vendor will tell you their agent passes. The useful questions are who was calling, what "pass" meant, and who checked the checker. Agent-to-agent evaluation is the design that lets you answer all three, and it is cheap enough that you can ask any vendor to run your own scenarios in front of you.
The older approach records caller audio, plays it at the agent, and waits for silence to decide the agent has finished. It is fragile because silence thresholds differ by deployment, it cannot judge whether the answer was any good, and it never exercises the real speech-to-text and text-to-speech path as a caller would. A synthetic caller that is itself a voice agent speaks each line as audio, knows exactly when it has finished speaking, and hears the target through the same client a person would use.
Buyers care because an evaluation built this way produces evidence you can inspect: transcripts, tool calls, per-turn scores with reasons, and latency measured where the caller hears it.
The mechanism
The research team's evaluation harness is one implementation of a general design. The pieces below are what to look for in anyone's.
A synthetic caller on the production path. The evaluator is a voice agent of its own. For each scenario turn it is handed the text to say and speaks it as audio into the target; its own audio-complete event marks the end of the evaluator's turn, so there is no silence detection and no guessing. The target runs in a real client (a scripted browser page or a real phone call), and an interceptor on the client captures the target's transcripts, function calls and speaking state. Because the evaluator speaks through text injection rather than a microphone, there is no echo path to confuse either side.
Personas and scenarios. A persona is how a caller behaves: patient or impatient, technical or not, cooperative or hostile, plus a voice style. A scenario is a persona plus an ordered list of turns and a description of what it tests. The research team's harness shipped six persona archetypes and thirty scenarios; the archetypes that earn their place are the hostile ones (the frustrated caller who pushes back, the caller comparing you with a rival, the caller who asks vague questions and expects the agent to lead). Scenarios should be generated and edited by people who know the calls, and the buyer should be able to add their own in minutes.
A rubric-scored judge. Every target response is scored by a language model against a fixed rubric with small integer scales and concrete fail cases.
| Dimension | Pass (2) | Partial (1) | Fail (0) |
|---|---|---|---|
| Relevant | Directly answers or performs the action | Partly on point | Misses the question |
| Grounded | Nothing invented | Minor uncertainty | Invented facts |
| Spoken prose | Plain sentences suitable for speech | Slightly verbose | Contains markup, lists or code formatting |
| On task | Stays within the agent's job or a valid hand-off | Slight tangent | Off topic or refuses without cause |
Four dimensions at 0 to 2 give a score out of 8; the research team's harness used 6 of 8 as the pass line. The judge runs through a serial queue (one call in flight, retries with backoff), and when every retry fails the turn receives a neutral score that must be flagged in the report, because unflagged neutral scores quietly inflate pass rates.
A latency breakdown per turn. Time to first audio measured at the client from the end of the evaluator's audio to the target's first playable audio; and pipeline stage times (speech-to-text to language model, language model to text-to-speech) from the events the client intercepted. Judge time is never in the caller's latency. Report median and worst per scenario, and per turn type: turns with a tool call and turns without.
A report you can replay. One file per run with metadata (target, timestamp, persona, model configuration), one entry per scenario with pass and duration, and one entry per turn with the caller text, the agent's response, the scores, the rationale, flagged issues and the latency triple. A dashboard is a convenience; the file is the evidence. Cross-run comparison by configuration and persona is where regressions show.
Judge limitations, stated. A judge that does not know the agent can navigate a page will fail every "I have opened the pricing page for you" as a hallucination. A judge shares the biases of its model family and may favour longer answers. A judge cannot hear audio quality: stutter, a clipped word, a mispronounced name. So: give the judge the agent's real capabilities; keep fail cases concrete; sample a fixed share of turns per run (a tenth to a fifth, plus every fail) for human review; record the disagreement rate; and keep a human listening to the audio of at least one scenario per run.
Evidence
The research team's evaluation harness in 2026 is the source for the design. Method: a synthetic caller drove a target voice agent embedded in a web page, the client intercepted transcripts and function calls, a rubric judge scored every turn, and reports were reviewed by hand while the harness was being tuned.
- Turn boundaries from the evaluator's own audio-complete event removed the silence-threshold tuning that pre-recorded playback needed, and removed the false "agent finished" verdicts that silence detection produced on pauses inside an answer.
- The judge produced false fails on navigation turns until its instructions stated that the agent could actually open pages. The fix was context, not a threshold.
- A serial judge queue with retries and a minimum gap between calls eliminated bursts that had caused rate-limit failures; the neutral fallback score was retained but surfaced in the report so it could be excluded.
- Six personas and thirty scenarios were enough to expose regressions between configurations; the hostile personas produced most of the failing turns.
- The rubric's most useful fail case was mechanical: markup characters in spoken text. The agent that passes it has a speech-aware output layer; the one that fails is reading its own formatting aloud, a cousin of the agent reads its instructions aloud.
These are design outcomes, not benchmark scores, and none of them compares vendors.
How to test for it in a demo
- Bring five scenarios. Write them in your own words: a normal booking, an ambiguous request, a two-fact question, a caller who changes their mind, and one hostile caller. Ask the vendor to run them through their evaluation harness while you watch. A vendor with a harness will; a vendor without one will offer a demo instead.
- Read one report. Ask for the file from the run, not the dashboard. Check that each turn has the caller text, the response, the scores, a rationale and latency. Check for flagged neutral scores.
- Time one turn yourself. Take a turn from the report and measure it from the recording. If the report says 800 ms and your stopwatch says 1.6 s, ask where the report's number was measured.
- Ask for the rubric and the spot-check. What the judge scores, the pass line, how many turns were reviewed by a person last month, and how often the person disagreed. No spot-check means the pass rate is the judge's opinion.
- Ask what the judge cannot see. The honest answer includes audio quality, and describes a human listening rung. See the verification ladder.
- Change a persona. Make the hostile caller ruder and rerun. Watch whether the agent's on-task and grounded scores hold.
The evaluation is only as good as its weakest rung; a harness that never plays audio is the harness-only testing anti-pattern with a nicer dashboard.
Questions to ask vendors
The three frontmatter questions: how the agent is evaluated end to end, what the judge scores and how often people disagree with it, and where each latency number is measured. A vendor who answers all three with a report file in hand has an evaluation practice. One who answers with a pass rate has a number.
Questions to ask vendors
- 01
How do you evaluate the agent end to end: recorded audio with silence detection, or a synthetic caller that speaks and listens through the production path?
A good answer: A synthetic caller driven by scenarios and personas, turn boundaries from the evaluator's own audio-complete event, transcripts and tool calls captured from the real client.
- 02
What does your judge score, what is the pass threshold, and how often do humans disagree with it?
A good answer: Named dimensions with concrete fail cases, a stated threshold, a spot-check rate, and the disagreement rate from the last review.
- 03
In your latency reports, where is each number measured and what is excluded?
A good answer: Time to first audio measured at the client from the end of caller audio, per-stage breakdown from pipeline events, judge time excluded, and any neutral-fallback scores flagged.
Related
- Anti-pattern
- Anti-pattern
- Anti-pattern