Skip to content
Voice AgentBible

Fewer tool calls per turn: fold what the next step needs into the last result

Every tool the model chains inside one turn adds about a second. How to fold data into the previous result, when not to, and how to hear the difference in a demo.

By · 4 min read

Last verified 30 Sept 2026v1.0Published 30 Sept 2026

Best practiceLatencyimpact high

Each extra tool the model chains inside one turn is another language-model round trip of roughly a second; return the data the next step always needs inside the previous tool's result.

Why buyers care

The latency budget explains why a turn that calls a tool costs two model calls. This article is about the turn that calls two tools, because that is where demos go quiet.

A caller asks whether a procedure is covered. The agent verifies who the caller is, which is one tool. Then it checks eligibility, which is another. Each tool is followed by a model call to decide what to do with the result. The caller hears a pause, perhaps a filler, another pause, and then the answer. In the research team's demo traces each chained round trip cost roughly a second on top of the turn.

Buyers care because these are exactly the turns the agent exists for. A greeting is fast on every platform. The turn that answers the caller's real question is the one that decides whether the call felt like a conversation or a queue, and its cost is set by how the tools were designed, not by which voice was bought.

The mechanism

A tool turn has a fixed shape: the model decides to call a tool, the tool runs, the model reads the result and speaks. Chaining adds a loop.

Turn shapeModel callsTool callsApproximate cost in the research team's traces
Plain answer10One model call, roughly 0.7 s
Single tool21Roughly 1.4 s plus your system's response time
Two chained tools32Roughly 2.1 s plus two system responses
Three chained tools43Roughly 2.8 s plus three system responses

The fix is to look at which tool is always called after which, and fold.

Fold the data the next step always needs into the previous result. If every eligibility check follows an identity lookup on the same call type, the identity lookup returns eligibility. The model gets both facts in one result and speaks. There is no second tool to call, so there is no third model call.

Tell the model not to call the second tool. The fold only saves time if the model knows the data is already there. The description of the identity tool says it returns eligibility for this call type; the description of the eligibility tool says not to call it when the identity result already carries eligibility. The model reads those descriptions at the moment of choosing, which is where the instruction has to be.

Carry state so nothing is fetched twice. After every tool that opens a gate (identity verified, record found, slot booked), the server refreshes a short status line into the prompt. The model sees the state directly and does not re-fetch to be sure. When the platform drops a tool result, which happens, the status line is what keeps the call on track.

Return the cached result for duplicates. A second identical request within a few seconds returns the first result instead of running again. This protects against a model that retries, and it protects the system of record from double writes.

Do not fold everything. Folding grows result payloads, and large results cost tokens and time on the next model call. Fold what the next step always needs. Leave optional lookups (a price for a procedure the caller might not ask about, a history the agent rarely reads) as separate tools. The test is "always", not "sometimes".

Move your test rules when you move your tools. If a tool starts opening a gate as a side effect, the harness rule that says "no gated tool before its gate opened" is now wrong. In the research team's builds three clean runs were rejected by a stale rule after a fold. The fold and the rule change belong in the same commit.

Evidence

The research team's demo-build measurements are one set of traces on one flow, not a benchmark. Method: a scripted caller ran an authorisation flow that required identity verification and an eligibility check, with the same speech and language models, in a real browser; voice-to-voice latency was measured at the client and the median taken over the tool-backed turns of a session.

  • Before the fold, the flow's tool-backed turns measured roughly 2.3 to 2.5 s median. The turn shape was identity tool, model call, eligibility tool, model call.
  • After the identity result carried eligibility for that call type, and the tool descriptions said so, the same turns measured in the same class as a single-tool flow on the same build, roughly 1.5 to 2.0 s. The saving was close to one model round trip.
  • The model still occasionally called the eligibility tool when the description was vague. Naming the condition in the description ("do not call when the identity result includes eligibility") removed the extra call in subsequent runs.
  • A status line in the prompt that described state for a tool the call type did not own caused the model to demand information it should not have asked for. State lines must be scoped to the tools the call type actually has.

The numbers are specific to that build. The pattern is not: each chained round trip costs a model call, and a model call is a large fraction of the whole budget.

How to test for it in a demo

  1. Ask a two-fact question. "Am I covered for a crown, and what would I pay?" or "Is the earliest slot with the hygienist also one where a loan car is available?" Count the pauses. One pause and an answer means folded or well-batched. Two pauses, or a pause, a filler and another pause, means chained.
  2. Ask for the trace. Any team that has worked on latency has a per-turn trace showing tool calls and model calls. Ask for one turn's trace of the question you just asked. Two tool calls in one turn should come with a reason.
  3. Repeat a fact. Give your identifier, then two turns later ask something that needs it. If the agent asks for it again or visibly re-looks-up, state is not carried.
  4. Ask for two things at once early in the call. Callers do this constantly. A well-designed agent answers both in one turn or explicitly takes them one at a time. A chained design goes quiet.
  5. Listen for fillers. A filler placed to cover a chained wait is a symptom being treated. It is fine if it does not talk over you; see spoken fillers that talk over callers. It is not a fix.

Record the turn and time it against the latency budget. A two-fact turn over about 2.5 s on a good network is almost always a chain.

Questions to ask vendors

The three frontmatter questions: how many calls the slowest common turn makes and what was done about it, how the model knows not to call a tool whose data it already has, and what a lookup tool returns. A vendor who has done this work will name a specific fold and its saving. A vendor who has not will talk about the voice.

Questions to ask vendors

  1. 01

    On your slowest common turn, how many tool calls and model calls happen, and what have you done to reduce them?

    A good answer: A named turn, a count, and a specific fold: for example, the identity lookup returns eligibility so the model never calls a second tool.

  2. 02

    How does the model know not to call a tool whose data it already has?

    A good answer: Tool descriptions say so explicitly, the server refreshes state into the prompt after each tool, and duplicate calls within a few seconds return the cached result.

  3. 03

    What do you return from a lookup tool: only what was asked, or everything the next step will always need?

    A good answer: Everything the next step always needs, kept short, with optional data left to separate tools so results do not bloat.