Skip to content
Voice AgentBible

Playbook

What separates a voice agent that survives real callers from one that survives a demo. Each article names the mechanism, the evidence, and a probe you can run in a vendor demo.

Best practices

  1. Evaluationimpact medium
    Agent-to-agent evaluation: synthetic callers, personas, a rubric judge and honest latency

    Evaluate with a synthetic caller on the production path, personas including hostile callers, a rubric judge spot-checked by people, and latency measured at the client.

  2. Turn-takingimpact medium
    Barge-in done right: gate interruption on microphone energy and know your echo path

    Barge-in must require real caller energy at the microphone over a recent window, not just a speech-start event; the agent's own voice and line noise both trigger false interruptions.

  3. Safetyimpact high
    Confirmation guards: the yes must be latest, plain and specific to the action

    Before any consequential action, the guard must find a recent, plain, action-specific yes in the caller's latest transcribed utterance, never trusting the model's claim that consent was given.

  4. Promptingimpact high
    Design tools, not a giant prompt: gates in code, confirmations as evidence

    Rules that must always hold belong in tool code and gate checks, not in prompt text; the prompt decides what to say next, the code decides what may happen.

  5. Latencyimpact high
    Fewer tool calls per turn: fold what the next step needs into the last result

    Each extra tool the model chains inside one turn is another language-model round trip of roughly a second; return the data the next step always needs inside the previous tool's result.

  6. Operationsimpact medium
    Stall watchdogs: detect a silent agent or a lost transcript and recover once

    A hosted voice agent can go silent for a minute or more mid-call; a well-armed watchdog detects it in seconds and recovers exactly once, without repeating any action the caller already triggered.

  7. Latencyimpact high
    The latency budget: why about 800 ms voice-to-voice is the line

    Voice-to-voice latency is end-of-turn detection plus one or two model calls plus time to first audio; under about 800 ms feels conversational, over about 1.2 s feels like an IVR.

  8. Evaluationimpact high
    The verification ladder: from unit tests to a live call on the buyer's network

    Unit tests and scripted harness runs cannot hear audio; a voice agent is verified only after a real browser or phone run, golden recordings, and one live call from a network like the audience's.

Anti-patterns

  1. Turn-takingseverity medium
    Fillers that talk over the caller, and silently disarm your stall detection

    An injected filler is a second speech stream; it collides with the caller, confuses reply completion, and if it counts as a reply it hides the stall it was meant to cover.

  2. Turn-takingseverity high
    Fixed silence timeouts: the endpointing setting that makes a voice agent feel like an IVR

    A fixed silence timeout cannot tell a pause from an ending; it splits turns on mid-sentence pauses, holds turns open on trailing digits and noise, and adds a fixed delay to every reply.

  3. Evaluationseverity high
    Harness-only testing: your test rig cannot hear audio

    A text-and-timing harness cannot observe playback stutter, overlap or network jitter; the only evidence of a working voice agent is a real client on a realistic network with counters.

  4. Safetyseverity high
    Prompt-only compliance: 'never ask for a card number' in the prompt is not a control

    A rule that lives only in the prompt is a preference the model usually honours; a control is code that refuses regardless of what the model decides.

  5. Pricingseverity medium
    The advertised per-minute price: why a $0.06 headline becomes $0.10 to $0.33 all-in

    The advertised per-minute price is usually the orchestration fee alone; speech, language model, voice, telephony, numbers and tooling roughly double to quintuple it.

  6. Promptingseverity medium
    The agent reads its instructions aloud: system notes leaking through tool results

    Anything placed in a tool result is candidate speech; internal notes must be named as not-for-speech, kept terse, and backed by a persona rule, or the agent will read them to the caller.

  7. Evaluationseverity high
    Vendor-run demos: scripted happy paths, cached clips and no adversarial persona

    A demo the vendor drives is a rehearsed recording of the agent at its best; only a demo the buyer drives, with an adversarial persona, measures the agent.

  8. Evaluationseverity high
    Words from nothing: speech recognition fabricating text on silence and noise

    Every speech-recognition system tested fabricates words on speech-free audio at some rate; a word whose timestamps fall outside the audio is fabricated by construction.