Skip to content
Voice AgentBible

RFP checklist

29 questions in eight weighted sections. Each one says why it matters and what a good answer sounds like, so a weak answer is obvious in the room.

Reviewed 30 Sept 2026. Must-have questions are marked; a poor answer to one of those ends the evaluation.

1Architecture and control

weight 15%
  1. Which components (speech-to-text, language model, text-to-speech, telephony) are yours, and which are third-party? Can each be swapped? must-have

    Why it mattersLock-in and outage exposure live in the layers you cannot see.

    Good answerA named component list, with at least the language model and voice swappable by configuration, and a documented failover path.

  2. Where do business rules live: in the prompt, in code, or in a flow builder? How are they versioned and tested? must-have

    Why it mattersPrompt-only rules drift and cannot be audited.

    Good answerDeterministic gates (identity, consent, payment) enforced in code or tools; prompts versioned with test suites and rollbacks.

  3. How does the agent read from and write back to our systems of record (CRM, EHR, PMS, DMS, ticketing)? must-have

    Why it mattersA demo that writes to the vendor's dashboard proves nothing about your workflow.

    Good answerNamed connectors or an API pattern with a write-back demonstration into a sandbox of your actual system.

  4. What happens when a tool call fails or times out mid-conversation?

    Why it mattersReal backends are slow; callers hear the gap.

    Good answerExplicit timeout budgets, a spoken fallback, retry policy and a logged incident, with an example transcript.

  5. Can we run parts of the stack in our own cloud or on premises?

    Why it mattersData residency and regulator expectations in some markets require it.

    Good answerA clear yes/no per component with reference customers or a documented deployment guide.

2Latency, turn-taking and quality

weight 20%
  1. What is your median and 90th-percentile voice-to-voice latency, measured how, on which network, and can we reproduce the measurement? must-have

    Why it mattersUnder about 800 ms feels conversational; above about 1.2 s feels like an IVR.

    Good answerNumbers with the measurement method (from end of caller speech to first agent audio), a way to run the same test on your phone, and honesty about tool-call turns being slower.

  2. How does the agent detect end of turn, and how does it behave when the caller pauses mid-sentence, speaks over it, or stays silent for eight seconds? must-have

    Why it mattersThese three moments decide whether callers trust the agent.

    Good answerA description of end-of-turn detection beyond a fixed silence timer, configurable barge-in, and a graceful silence prompt, all demonstrable live.

  3. How do you handle accents, code-switching between languages, background noise and telephone-quality audio?

    Why it mattersYour callers do not sound like the demo audio.

    Good answerNamed language and accent coverage, a test on your own recorded calls, and known limits stated plainly.

  4. How does the agent confirm numbers, dates, names and spellings before acting on them?

    Why it mattersMisheard digits become wrong bookings and wrong payments.

    Good answerRead-back for high-stakes fields, configurable per field, with examples of recovery from a misrecognition.

  5. Can we listen to unedited recordings of production calls (with consent) rather than curated demos?

    Why it mattersCurated clips hide the failure modes.

    Good answerYes, with a sample of ordinary calls including at least one that went badly and how it was handled.

3Integrations and telephony

weight 10%
  1. Which contact-centre platforms and carriers can you connect to (SIP, existing numbers, CCaaS), and what does porting or forwarding our numbers involve? must-have

    Why it mattersTelephony is where projects stall.

    Good answerNamed integrations, a porting or forwarding plan, and who owns the numbers.

  2. How do warm transfers to a human work, including context hand-off and what the human sees? must-have

    Why it mattersTransfers without context restart the conversation.

    Good answerA transfer with transcript, intent and collected fields visible to the agent, demonstrable live.

  3. Which calendar, CRM and scheduling systems do you support natively, and how are conflicts and time zones handled?

    Why it mattersDouble bookings are the most common early failure.

    Good answerNamed systems and a demonstration of conflict handling.

4Compliance and consent

weight 15%
  1. How do you handle consent, do-not-call scrubbing and calling-hour restrictions per jurisdiction we operate in? must-have

    Why it mattersAutomated outbound calls are regulated almost everywhere.

    Good answerPer-jurisdiction configuration, suppression-list integration, an audit log of consent, and named rules (for example TCPA, PECR, TCCCPR).

  2. Does the agent disclose that it is AI, and can we control the wording and timing? must-have

    Why it mattersDisclosure is legally required in some markets and expected in most.

    Good answerConfigurable disclosure at the start of the call, logged per call.

  3. How are call recordings and transcripts stored, for how long, where, and who can access them? must-have

    Why it mattersRecording consent and retention rules vary by country and sector.

    Good answerRegion-pinned storage, configurable retention, role-based access, and deletion on request.

  4. Which sector frameworks do you support (for example HIPAA business associate agreements, PCI scope reduction, FDCPA call-frequency limits)?

    Why it mattersSector rules are where fines actually land.

    Good answerSigned agreements available, PCI pause-and-resume or tokenised capture, and rule enforcement in code rather than prompt.

5Security and data handling

weight 10%
  1. Which independent certifications or audits do you hold, and can we see the reports?

    Why it mattersClaims are cheap; reports are not.

    Good answerCurrent SOC 2 Type II or ISO 27001 report shared under NDA, with scope covering the voice stack.

  2. Is our call data used to train or improve models, and can we opt out? must-have

    Why it mattersCustomer audio is sensitive and often contractually protected.

    Good answerA contractual no-training default or an explicit opt-out that applies to every sub-processor.

  3. How do you prevent prompt injection and social engineering of the agent into revealing or changing data?

    Why it mattersCallers will try.

    Good answerTool-level authorisation independent of the model, red-team results, and a description of what the agent can never do regardless of what it is told.

6Pricing and total cost

weight 10%
  1. Provide the all-in cost per connected minute including platform, speech-to-text, language model, text-to-speech, telephony and number rental, for our call profile. must-have

    Why it mattersAdvertised base rates commonly understate the total by two to four times.

    Good answerA line-item breakdown for your volumes, with what changes at 2x and 5x volume.

  2. What one-off and recurring costs exist beyond minutes: integration, prompt and flow maintenance, testing, support tiers, minimum commitments?

    Why it mattersMinutes are rarely the largest line in year one.

    Good answerA transparent list with typical ranges from comparable customers.

  3. How are transferred, abandoned, silent and failed calls billed?

    Why it mattersBilling edge cases quietly inflate invoices.

    Good answerClear rules, ideally no charge for calls the agent failed to answer.

7Testing, monitoring and operations

weight 10%
  1. How do we test the agent before and after each change: scripted scenarios, synthetic callers, regression suites? must-have

    Why it mattersEvery prompt change can break a working flow.

    Good answerA test harness we can run, with scenario coverage reports and release gates.

  2. What do we see in production: per-call transcripts, latency breakdown, tool outcomes, containment, transfer reasons?

    Why it mattersYou cannot manage what the dashboard hides.

    Good answerPer-call traces with timings, exportable, plus alerts on latency and failure spikes.

  3. What is your incident history over the last twelve months and how were customers informed?

    Why it mattersEvery stack has outages; the question is transparency.

    Good answerA public or shared status history with post-incident reports.

8Commercial terms and exit

weight 10%
  1. What are the contract length, termination terms, and data-export format at exit? must-have

    Why it mattersYou should be able to leave with your transcripts, prompts and configurations.

    Good answerMonth-to-month or annual with a clean exit; full export of call data and configurations in open formats.

  2. Who owns the prompts, flows and voice configurations we build on your platform?

    Why it mattersIntellectual-property clauses decide how portable your work is.

    Good answerThe customer owns everything they author.

  3. What service levels do you commit to for availability, latency and support response, and what are the remedies?

    Why it mattersAn SLA without remedies is a wish.

    Good answerNumeric commitments with credits or termination rights.