Skip to content
Voice AgentBible

Methodology

How Voice Agent Bible defines and measures latency, containment and turn-taking, scores demos, probes speech fabrication, models cost and sources compliance.

By · 5 min read

Last verified 30 Sept 2026v1.0Published 30 Sept 2026

This page describes how the site defines its terms, how it measures what it labels Measured, and the assumptions behind its models. Readers should be able to reproduce every method with a phone, a recorder and a spreadsheet.

Definitions

Time to first audio (TTFA). The delay between a text-to-speech request starting and the first audio sample being available to play. A component metric. We report it only when the measurement point is stated.

Voice-to-voice latency. The time from the end of the caller's speech to the first audio of the agent's reply, as heard by the caller. This is the end-to-end figure and the one we use for all comparisons. It includes end-of-turn detection, final transcription, model response, any tool call, synthesis start-up and the network path back to the phone.

End-of-turn detection. The system's decision that the caller has finished speaking. We measure it indirectly through voice-to-voice latency on turns where the caller's speech ends cleanly, and directly through false cut-off counts on turns where the caller pauses mid-utterance.

Containment. A call counts as contained only when the caller's intent was completed and the result exists in the system of record: an appointment in the schedule, a status delivered against the correct record, a payment posted. Calls that ended without a transfer but without an outcome are not contained; they are counted as abandoned or unresolved.

How latency is measured

Calls are placed from a real browser or a real phone, never from a vendor's test console. The channel is stated with every result: mobile handset on a named carrier type, landline, or browser over a described connection. Network conditions are stated: home broadband, office network, mobile data, and whether any throttling or packet loss was applied.

Timing is taken from the call recording. The end of caller speech is marked at the last voiced sample; the start of agent audio is marked at the first agent sample above the noise floor. Marking is done in an audio editor at millisecond resolution and checked by a second pass.

Sample size is stated with every result and is never fewer than twenty turns per condition. We report the median and the 90th percentile; we do not report means, because a few slow turns dominate them and hide the caller experience. Turns that call a tool (checking availability, looking up a record) are reported separately from turns that do not, because the two distributions are different and vendors tend to quote the faster one.

Demo scoring rubric

Each demo script on this site has a scoring sheet. Four areas are scored.

  1. Functional beats. Each step of the call (greeting and disclosure, intent capture, identity confirmation, the core task, the write to the system of record, the close) is marked complete, partial or missed.
  2. Latency. Median and 90th-percentile voice-to-voice latency across the call, with tool-backed turns separated, using the method above.
  3. Recovery. How the agent handles interruption, mid-sentence pauses, a misheard detail, a request for a human, silence, and an out-of-scope question. Each is described, not scored numerically.
  4. Compliance disclosures said aloud. Whether the agent identified itself as automated, announced recording where required, identified the business and gave a callback or contact route, unprompted.

Traps embedded in the script are scored pass or fail only, with the pass and fail conditions written before the call. A trap with an ambiguous result is scored fail and the ambiguity noted.

Agent-to-agent evaluation design

For volume testing, a synthetic caller, itself a voice agent, places real phone calls to the agent under test. It is configured with personas: at minimum three ordinary callers of varying pace and clarity and one adversarial caller who attempts to extract another person's data or to push the agent outside its scope. Each persona has a goal and a set of behaviours (interrupts, pauses, changes of mind, code-switching where relevant).

Transcripts and recordings are scored by a language-model judge against a written rubric matching the demo scoring sheet. The judge's limitations are treated as part of the method: it can be fooled by fluent wrong answers, it favours longer transcripts, and it cannot hear latency or tone. To control for this, a fixed share of judged calls, never less than one in ten and higher for any trap with a fail rate under five percent, is reviewed by a person, and the disagreement rate between judge and reviewer is published with the result. Latency from synthetic runs is reported separately from human-placed calls, because the synthetic caller adds its own delays.

Speech-recognition fabrication probe

Some speech recognisers produce words from audio that contains no speech. The probe is simple and reproducible. Audio clips of silence, of steady noise (line noise, comfort noise, room tone) and of non-speech filler (breathing, hold music, a cough) are prepared at several durations, typically from one to thirty seconds, and are submitted to the recogniser in the same streaming configuration an agent would use. Any words returned are fabrications by construction.

A second detector catches fabrications on real speech: word timestamps that lie beyond the length of the audio submitted. A recogniser that reports a word ending at twelve seconds in a ten-second clip has invented it. Each clip is submitted several times, because fabrication is often intermittent, and the rate is reported per condition. Results are published as rates per condition across unnamed systems, never as a named-vendor ranking.

TCO model assumptions

The all-in cost calculator uses list-price ranges taken from vendors' public pricing pages, each with a retrieval date. It builds an all-in cost per connected minute from layers: platform fee, speech recognition, language model, text-to-speech, telephony (inbound leg, outbound leg, transfer legs), a share of human handling driven by the transfer rate, and amortised integration and maintenance.

The human baseline is a loaded hourly cost adjusted for occupancy, after-hours premium and handle minutes per call, with defaults stated. Region multipliers adjust telephony and language costs for markets where they differ materially. Every output is an estimate built from public list prices and stated assumptions. It is not a quote, and actual contracts differ.

Compliance matrix sourcing

Every row of the compliance matrix is built from the primary text first: the statute, the regulation, the regulator's own guidance page. Secondary reports are used only to flag pending changes and are marked as Reported. Each row carries a confidence column (high, medium, low) reflecting how directly the primary text answers the question for a voice agent, and a verified date. Rows are re-verified quarterly. Items known to be in motion (rulemakings, amendments, effective dates that may shift) are listed in a watch field on the row and repeated in the Weekly regulatory watch. Nothing in the matrix is legal advice.

What the site does not claim

In its first year, the site publishes no vendor rankings or scores from its own numbers. Measured results appear as methods and ranges. The site does not claim to have tested any named vendor unless a Measured source on the page says so and describes the method. It does not claim analyst credentials, customer references or knowledge of any vendor's unpublished plans.

Change history

DateVersionChange
2026-09-301.0Initial publication of the methodology.