Vendor scorecard
Eight weighted sections. Score each vendor 0 to 5 per section from the evidence they gave you in the demo and the RFP, multiply by the weight, add up. The site does not publish vendor scores; this is your sheet.
| Section | Weight | Must-haves (a zero here eliminates) | Vendor A | Vendor B | Vendor C |
|---|---|---|---|---|---|
| Architecture and control | 15% | Which components (speech-to-text, language model, text-to-speech, telephony) are yours, and which are third-party? Can each be swapped? · Where do business rules live: in the prompt, in code, or in a flow builder? How are they versioned and tested? · How does the agent read from and write back to our systems of record (CRM, EHR, PMS, DMS, ticketing)? | __ × 15 | __ × 15 | __ × 15 |
| Latency, turn-taking and quality | 20% | What is your median and 90th-percentile voice-to-voice latency, measured how, on which network, and can we reproduce the measurement? · How does the agent detect end of turn, and how does it behave when the caller pauses mid-sentence, speaks over it, or stays silent for eight seconds? | __ × 20 | __ × 20 | __ × 20 |
| Integrations and telephony | 10% | Which contact-centre platforms and carriers can you connect to (SIP, existing numbers, CCaaS), and what does porting or forwarding our numbers involve? · How do warm transfers to a human work, including context hand-off and what the human sees? | __ × 10 | __ × 10 | __ × 10 |
| Compliance and consent | 15% | How do you handle consent, do-not-call scrubbing and calling-hour restrictions per jurisdiction we operate in? · Does the agent disclose that it is AI, and can we control the wording and timing? · How are call recordings and transcripts stored, for how long, where, and who can access them? | __ × 15 | __ × 15 | __ × 15 |
| Security and data handling | 10% | Is our call data used to train or improve models, and can we opt out? | __ × 10 | __ × 10 | __ × 10 |
| Pricing and total cost | 10% | Provide the all-in cost per connected minute including platform, speech-to-text, language model, text-to-speech, telephony and number rental, for our call profile. | __ × 10 | __ × 10 | __ × 10 |
| Testing, monitoring and operations | 10% | How do we test the agent before and after each change: scripted scenarios, synthetic callers, regression suites? | __ × 10 | __ × 10 | __ × 10 |
| Commercial terms and exit | 10% | What are the contract length, termination terms, and data-export format at exit? | __ × 10 | __ × 10 | __ × 10 |
| Weighted total (max 500) | 100% | ___ | ___ | ___ |
How to score honestly
- Score from evidence you saw, not from the vendor’s answer. A claim without a demonstration scores at most 2.
- Any must-have scored 0 eliminates the vendor regardless of total.
- Have two people score independently and reconcile; disagreements point at the questions to re-ask.
- Keep the sheet with the recordings and traces from the bake-off so the decision can be audited later.
Print this page (it is styled for print) or copy the table into a spreadsheet. Pair it with the demo scoring sheet, which scores the live run rather than the paperwork.