Skip to content
Voice AgentBible

Design tools, not a giant prompt: gates in code, confirmations as evidence

Why a 20,000-word prompt fails where a tool-based design does not: gates in code, confirmations as transcript evidence, and how to probe for it in a demo.

By · 5 min read

Last verified 30 Sept 2026v1.0Published 30 Sept 2026

Best practicePromptingimpact high

Rules that must always hold belong in tool code and gate checks, not in prompt text; the prompt decides what to say next, the code decides what may happen.

Why buyers care

Most voice agents you will be shown are, underneath, one long block of text handed to a language model on every turn. The text says who the agent is, what it may not say, how to handle an emergency call, how to read a phone number back, and what to do when the caller goes quiet. It grows every time something goes wrong in a call, because adding a paragraph is the only lever the builder has.

That lever stops working. In the research team's demo builds, a prompt-only agent for a booking line had grown past 20,000 words and roughly fifty headed sections before anyone asked why it still booked the wrong date. The prompt contained a full calendar table so the model could work out "next Friday" by reading. The same job, rebuilt around four tools and a prompt of a few hundred words, stopped making date errors because the model no longer computed dates at all. It asked a function.

Buyers care because the two designs fail differently. A prompt-only agent fails silently and randomly: the rule was there, the model did not follow it this time. A tool-based agent fails loudly and reproducibly: the gate refused, the log says why, and the fix is a line of code rather than another paragraph of pleading. If you are buying something that will touch a schedule, a payment or a patient record, you want the second kind.

The mechanism

A language model is good at deciding what to say next and poor at holding forty rules in mind while it does so. The design that works splits the job.

ConcernPrompt-only designTool-based design
Availability, prices, hoursWritten into the prompt, stale by next weekA function reads the source of record
Date arithmeticThe model reasons over a calendar tableA function returns day and date; the model reads it back
"Verify identity before disclosing"A sentence the model usually obeysA precheck in code refuses the disclosure tool until identity is verified
"Confirm before you submit"A sentence the model usually obeysThe submit tool requires a recent, plain, specific yes in the transcript
Duplicate submissions"Never submit twice"The same request within about 15 seconds returns the original result
State (verified, booked, recorded)The model infers it from the conversationThe server refreshes the prompt with a status line after every gate opens

Three practices carry most of the weight.

Gates live in code. Every tool that changes or discloses something has a precheck that runs before the model's request is honoured. The precheck looks at server-side state (is this caller verified, has the criteria list been completed, has this item already been recorded) and refuses with a short structured reason. The model then has to ask the missing question. It cannot talk its way past the gate, and the prompt does not need to describe the gate at all.

Confirmations are transcript evidence. When a tool needs the caller's consent, the guard does not trust the model's claim that consent was given. It reads the caller's latest transcribed utterance and checks three things: it came after the agent asked, it is a plain affirmative of a dozen words or fewer with no question in it, and it is specific to the action. The confirmation guards article covers the timing traps.

Tool descriptions carry behaviour the prompt cannot enforce. "Verify once per call." "Contact details are on file; never ask for a phone number." "Call this the moment the caller picks a slot; never write your own confirmation question." A description is read at exactly the moment the model is choosing that tool, which is when the instruction matters. The same sentence buried in a long prompt competes with everything else in it.

The prompt that remains is short: role, voice, a handful of facts about the location, a checklist for the booking flow, and one sentence that does most of the work: for any date, availability, price or hours answer, call a function and never compute it yourself.

Evidence

The research team's demo builds across several verticals in 2026 show the pattern rather than a benchmark. Method: the same booking or claims flow was run with scripted callers against a prompt-only design and a tool-based design on the same speech and language models, and failures were classified from recorded transcripts and server logs.

  • The prompt-only booking agent's prompt had grown to roughly 23,000 words in about 50 sections, several duplicated, with a calendar table for date lookup. Date errors and "asked a question the caller had already answered" were the two most common scripted-run failures.
  • The rebuilt agent used a prompt of roughly 450 words and four tools: check availability, answer location questions, arrange a callback, save a call summary. Date errors disappeared because dates came from a function; the model's only job was to read back the day and the date.
  • In a claims flow, a status line describing identity verification leaked onto a call type that had no identity tool the moment the prompt was refreshed, and the agent demanded verification before a simple booking. The fix was code, not prompt: status lines are emitted only when the call type owns the tool.
  • A tool result field named "note" was sometimes read aloud to the caller. Naming the field so its purpose is unmistakable, and keeping it terse, stopped the leak. See the agent reads its instructions aloud.
  • The hosted platform occasionally dropped a tool result mid-call. The refreshed status line in the prompt was the only state that survived, and the call completed.

None of these are vendor comparisons. They are the failure classes a buyer will see in any demo built the prompt-only way, whoever built it.

How to test for it in a demo

You can tell the two designs apart from the caller's chair.

  1. Ask for a date the model has to compute. "Next Thursday" on a Wednesday, or "the Monday after the holiday". A tool-based agent names a day and a date in one turn and gets it right. A prompt-only agent hesitates, gets it right sometimes, or names a day without a date.
  2. Try to skip a gate. Ask for something that should need identity verification before you have verified. A gated agent asks the missing question every time. A prompt-only agent sometimes lets it through, and you will not know which time you got.
  3. Say yes to nothing. Before the agent asks a confirmation question, say "yes, go ahead and book it." A guarded agent asks the question anyway. A prompt-only agent may write the booking on your instruction alone.
  4. Repeat yourself. Give the same instruction twice within a few seconds. Check the vendor's log or your sandbox for one write, not two.
  5. Ask about the prompt and the tool list. Vendors rarely show the prompt itself, but they should be able to say how long it is, how many tools exist, and which rules are enforced in code. "It is all in the prompt" is the answer you are testing for.
  6. Change a fact. Ask them to change opening hours or a price during the demo. If it takes a prompt edit and a redeploy, the facts live in the prompt and will drift in production.

Score each probe pass or fail and keep the transcript. Compare with the prompt-only compliance anti-pattern, which is the same problem applied to disclosure and consent rules.

Questions to ask vendors

The three questions in the frontmatter, in the order a weak answer should worry you: which rules are enforced outside the prompt, how a confirmation is verified, and how state survives a lost tool result. For regulated verticals add a fourth: "Show me the log line that proves the gate refused." A vendor who has built gates in code will have one, and will be pleased you asked.

Questions to ask vendors

  1. 01

    Which of the agent's rules are enforced outside the prompt, in code, and which rely on the model following instructions?

    A good answer: A short list of code-enforced gates (identity before disclosure, confirmation before submit, duplicate suppression, completeness checks) and an honest admission that tone and phrasing rely on the prompt.

  2. 02

    When the agent submits or books something, how does the system verify the caller actually agreed?

    A good answer: The guard reads the caller's latest transcribed utterance and checks it is a recent, plain, action-specific yes; the model's own claim of consent is not enough.

  3. 03

    How long is the system prompt, how many tools does the agent have, and what happens when the platform loses a tool result mid-call?

    A good answer: A prompt measured in hundreds of words, a handful of tools with behavioural descriptions, and server-side state that is refreshed into the prompt so a lost result does not lose the call.