Skip to content
Voice AgentBible

Golden recording

Golden recordings for voice-agent evaluation: consented real call audio with a checked transcript, used to test speech accuracy the same way across vendors.

By · 1 min read

Last verified 30 Sept 2026v1.0Published 30 Sept 2026

Glossary

A real call recording, consented and redacted where required, paired with a human-checked reference transcript. A set of them is the fixed yardstick for measuring speech accuracy and replaying the same caller behaviour against every vendor.

Also called: golden set, reference recording, test set, ground-truth audio, reference transcript, evaluation corpus.

What it is

A golden recording is audio you trust, with a transcript you trust. Typically it is a set of several dozen to a few hundred real calls drawn from your own lines: the accents, the background noise, the code-switching, the mumbled dates of birth and the hold music that your callers actually produce. Each has a reference transcript checked by a person, with a stated normalisation convention (how numbers, abbreviations and punctuation are written). Where privacy rules require it, personal data is redacted or the recordings are synthesised by staff following real scripts.

The set is frozen. Every vendor is measured against the same files, and the same files are re-run after each vendor update.

Why it matters when buying

Without golden recordings, speech accuracy claims cannot be compared, and you are choosing on datasheet numbers measured on someone else's audio. With them, you can compute word error rate and entity accuracy yourself, spot a regression after a model change, and replay hard calls (the emergency, the angry caller, the number read in groups) against each agent. They are also the cheapest evaluation asset to build: a week of recordings and a few days of transcription.

What to ask

Ask yourself first: do you have consent and a lawful basis to use these recordings for testing, and have you redacted what you must? Then ask vendors to run your set and return raw transcripts, not scores. Ask that the same set be re-run before each production model change. Ask how they handle the recordings afterwards (deletion, no training).

← All terms