Synthetic caller (agent-to-agent testing)
Synthetic callers for voice-agent testing: an AI that phones your agent with scripted personas at scale, what it can and cannot judge, and the human spot-check.
By Voice Agent Bible Research · 1 min read
Last verified 30 Sept 2026v1.0Published 30 Sept 2026
An automated caller, itself a voice agent, that phones the agent under test with scripted personas and goals, hundreds of times, so behaviour can be measured at scale. Scored by a rubric and a language-model judge, with human spot-checks because the judge is fallible.
Also called: agent-to-agent testing, simulated caller, AI caller, automated conversation testing, persona testing, voice agent simulation.
What it is
A synthetic caller is a second voice agent configured to act as a customer. It is given a persona (a hurried parent, a hard-of-hearing pensioner, a caller who code-switches, an adversarial caller trying to extract another patient's details), a goal, and a set of behaviours (interrupts, pauses mid-number, changes their mind). It places real calls over the phone network to the agent under test, many times, and records everything. A rubric and a language-model judge score each transcript on functional outcome, recovery and compliance statements; a person reviews a fixed share of the calls to check the judge.
Why it matters when buying
Human testers can run twenty calls; a synthetic caller can run two thousand, overnight, across every persona, and rerun them after each vendor change. That turns evaluation from anecdote into distribution: how often does the agent fail the emergency trap, not whether it failed once. The limits matter equally. The synthetic caller's own speech is cleaner and more regular than a human's, so it under-tests turn-taking and accent handling. The judge can be fooled by fluent but wrong answers and is biased toward verbose transcripts. Latency measured this way includes the synthetic caller's own delays and must be separated from the agent's.
What to ask
Ask a vendor offering simulation what share of judged calls are human-checked and how disagreements are handled. Ask whether personas include adversarial and impaired callers. Ask how the synthetic caller's speech timing is varied. Treat results as a screen that finds problems, not proof that there are none; golden recordings and live human calls stay in the bake-off.
Related
- Term
- Term
- Term
- Term
- Term
- Term
- See also
- See also