Drive-thru voice AI: noise, upsell rules, order accuracy and crew hand-off
What drive-thru voice AI must do in wind, engine noise and car audio: accurate orders, upsell by rule, crew hand-off, reported chain deployments and a lane test.
By Voice Agent Bible Research · 5 min read
Last verified 01 Oct 2026v1.0Published 01 Oct 2026
KPIs at a glance
| KPI | Typical baseline | Target | How to measure |
|---|---|---|---|
| Order accuracy at the window | Your current accuracy from mystery-shop or remake counts, measured for one month per lane before any pilot | No worse than the crew's measured accuracy in month one, better by month three, counted at the pickup window | Remakes and refunds per 100 orders by lane, plus a weekly audio sample compared with the POS order. |
| Intervention rate | Not applicable before deployment | Under 15% of orders need a crew member to take over the speaker (rule of thumb from reported deployments; menus with heavy customisation sit higher) | Orders where the crew took over / orders attempted by the agent, by lane and by hour. |
| Time at the speaker | Your drive-thru timer's current order-point time per lane | At or below the crew's median at the order point; no increase at the 90th percentile | Timer system order-point duration, agent orders against crew orders, same lane, same daypart. |
| Upsell within rules | Your current attachment rate on the items you want suggested | Attachment lift on the configured items, with zero suggestions outside the rule set (one offer, declined offers not repeated) | Attachment rate by item against baseline; transcript audit for repeated or off-rule offers. |
| Hand-off time | Not applicable | Under 3 seconds from the agent deciding it cannot proceed to a crew voice on the speaker, with the partial order on the crew screen | Timestamp of hand-off trigger to crew audio, from recordings, per hand-off. |
| Voice-to-voice latency in the lane | Rule of thumb used across this site: above about 1.2 s per turn the agent feels like an IVR; in a lane it also feels like the speaker is broken | Median under 0.7 s; 90th percentile under 1.3 s on turns that read the menu or update the order | End of customer speech to first agent audio from lane recordings, tool-backed turns only. |
What it is
A drive-thru voice agent greets the car at the speaker post, takes the order with modifiers and quantities, offers the upsell your rules allow, reads the order back, writes it to the POS so it appears on the kitchen and window screens, and tells the customer to pull forward. When it cannot proceed, because the audio is unusable, the request is outside the menu or the customer asks for a person, it hands off to the crew headset in one move with the partial order intact.
Everything about this channel is harder than the phone. The microphone is outdoors. The customer has an engine running, a window half down, wind, a radio and passengers. The agent's own voice comes back through the speaker post into the microphone. Orders are fast, slangy and full of mid-order changes. And the timer is running: a lane that slows by ten seconds per car loses cars.
Public press reports large deployments in the United States. Restaurant Dive reported in July 2026 that one chain had voice AI in more than 890 restaurants across 38 states, after its parent announced in a March 2025 filing a rollout to 500 restaurants across four brands, and after the same parent slowed the rollout in August 2025 following customer complaints and prank orders. The same publication reported in May 2025 that another chain had the technology in over 160 restaurants with a goal of more than 500 by the end of the year, and in January 2026 that a third chain's trial with its technology partner had ended. Read those as reported, not as proof that any product will work on your lane.
Who buys it
- Quick-service chains and large franchisees with drive-thru as the majority of sales, who need lane-by-lane accuracy and timer data, central menu and upsell rules, and a procurement process that includes the headset and timer vendors.
- Regional chains and multi-unit franchisees piloting in a handful of stores before a system-wide decision, usually driven by labour availability at the order point.
- Drive-thru hardware and POS vendors' customers choosing between a bundled agent and a standalone one on the same lane test.
Budget owner: the chief operating officer or head of restaurant technology at chain level; the franchisee principal at store level, with the brand's technology standards team approving the integration.
KPIs
Measure the crew first. One month of order accuracy from remakes and refunds, timer order-point duration, and attachment rate on the items you care about, by lane and daypart. The agent is judged against that, not against a vendor slide. Then track accuracy at the window, intervention rate, time at the speaker, upsell within rules, hand-off time and voice-to-voice latency in the lane.
Two traps. Accuracy at the speaker is not accuracy at the window; count remakes. And an intervention rate that falls because the agent stopped handing off is an accuracy problem in disguise. Read the two together, by hour, because the dinner rush and the late-night lane behave differently.
Demo script
Insist on a lane test at one of your stores, on your microphone, speaker, headset and POS, with your menu and your upsell rules loaded. Use a real car.
- Greeting under engine noise. Pull up with the engine running and the window half down. Pass: a short greeting, an AI disclosure if your brand requires one, and an open question. Fail: a long scripted welcome or a greeting that triggers again on the engine.
- Fast order with modifiers. "Two number threes, one no onions, one with extra pickles, a large drink and a kids' meal with apple slices." Pass: each modifier on the right item, read back in order.
- Quantity change within ten seconds. "Actually make that three number threes." Pass: one line updated, no duplicate.
- Background noise trap. Turn the radio up and have a passenger talk for twenty seconds without ordering. Pass: the agent waits and adds nothing. Fail: it adds an item, or says "I didn't catch that" repeatedly.
- Interruption over the read-back. Talk over the agent's read-back with "no, no pickles on the second one". Pass: it stops within a word and fixes the modifier. Fail: it finishes the read-back or treats its own echo as speech.
- Upsell by rule. Decline the first suggestion. Pass: one offer, remembered as declined, no second attempt. Fail: a repeated or off-rule offer.
- Ambiguous item. "A chicken sandwich" where you sell three. Pass: it asks which, with the names. Fail: it picks one.
- Eight seconds of silence after the agent asks "anything else?" Pass: one short prompt, then a read-back and total. Fail: a hang-up or a repeated greeting.
- Out-of-scope request. Ask for an item from another brand, then complain about the last visit. Pass: a plain "we don't have that" and a short hand-off offer for the complaint. Fail: an invented item or an apology loop.
- Hand-off. Mumble a long custom order the agent cannot parse. Pass: a crew voice on the speaker within a few seconds with the partial order on the crew screen, and the customer is not asked to start over. Then watch the order on the POS and the window screen.
Score each trap pass or fail. Added items during the noise trap, a repeated upsell, or a hand-off that loses the partial order are hard stops. A vendor who wants to demo from a recording in their lab has not passed the demo.
Compliance notes
A customer at the speaker post initiates the interaction, so the TCPA consent rule in the United States is not in play unless the deployment adds outbound contacts, in which case prior express consent and the 8 a.m. to 9 p.m. window apply. Recording is the live question: many drive-thru audio systems record for quality, roughly a dozen states require all-party consent, and a sign at the order point is the usual notice; your counsel decides whether the agent should also say it. Several states have bot-disclosure laws, so disclose by default. In the United Kingdom, recording the lane is personal-data processing under UK GDPR and needs notice and a documented lawful basis; the 14-allergen information duty applies to drive-thru food as to any other, so the agent answers allergen questions from data and does not guarantee. In the European Union, the AI Act's Article 50 duty to tell people they are talking to an AI applies from 2 August 2026, which at a speaker post means a disclosure in the greeting. In Australia, state surveillance-devices laws differ, so post and announce recording. Informational, not legal advice; the compliance rows on this page carry the sources.
Build or buy
Buy. This is the one restaurant use case where building is rarely sensible: the acoustic front end, the echo path through the speaker post, the headset integration and the lane-timer plumbing are specialist work, and the public record shows that even well-funded chains have paused or ended rollouts. Your leverage is the pilot design: your lanes, your menu, your crew's measured accuracy as the baseline, a noise trap in every demo, and an exit clause tied to accuracy at the window and intervention rate by month three.
Questions to ask vendors
- 01
Run your agent on our lane, with our microphone and speaker, with a car idling and a radio on, and show me the order on our POS.
A good answer: A lane test on your hardware, not a recording from the vendor's lab. The order appears on the POS and the crew screen during the test with modifiers correct.
- 02
What happens when the car's engine, the wind or a passenger produces speech-like noise while nobody is ordering?
A good answer: The agent waits; it does not add items or say 'sorry, I didn't catch that' every four seconds. A word whose timestamps fall outside real speech is treated as fabricated.
- 03
How does the agent hand off to the crew, and what does the crew see when it does?
A good answer: A crew voice on the speaker within a few seconds, with the partial order and the reason on the headset or screen. The customer is not asked to start again.
- 04
What are the upsell rules, who sets them, and how do you stop the agent repeating an offer the customer declined?
A good answer: A rule set you configure per item and daypart, one offer per order, declines remembered within the order, and an audit of offers made.
- 05
How does the agent handle 'actually make that two' and a customer who talks over the read-back?
A good answer: One line updated, no duplicate, and the agent stops speaking when the customer does because barge-in is gated on real speech energy, not on its own echo from the speaker post.
- 06
What does the agent do with a request it cannot fulfil, like an item from another brand or a complaint about the last visit?
A good answer: A short, plain answer and a hand-off if needed, not an invented item or an apology loop.
- 07
What accuracy and intervention rates have you measured at the window in live stores, and how were they measured?
A good answer: Numbers from the pickup window with the method stated, not accuracy at the speaker. If the vendor cites press, you will verify it yourself.
Matrix rows that apply
Rows from the global compliance matrix that apply to this page. Informational only, not legal advice; dates change, confirm with counsel and the regulator.
| Jurisdiction | Consent for automated calls | AI disclosure | Calling hours | Recording | Verified |
|---|---|---|---|---|---|
| United States (federal)confidence high | Required The FCC's February 2024 declaratory ruling confirms that AI-generated or cloned voices are "artificial or prerecorded" voices under the TCPA. Outbound calls using them need prior express consent; marketing calls to mobile numbers need prior express written consent. Inbound calls initiated by the consumer are outside this consent rule. | Conditional No federal statute yet requires an agent to announce that it is AI. TCPA rules already require prerecorded or artificial-voice calls to identify the caller at the start and give a callback number. An FCC proposal (2024) would add an explicit AI disclosure; several states have their own bot-disclosure laws. Disclose by default. | Required Telephone solicitations only between 8 a.m. and 9 p.m. in the called party's local time (47 CFR 64.1200(c)(1)). | Conditional Federal law is one-party consent; roughly a dozen states (including California, Florida, Washington and Pennsylvania) require all-party consent. Announce recording at the start of every call unless counsel confirms otherwise. | 2026-09-30 |
| United Kingdomconfidence medium | Required The ICO treats conversational AI voice calls as automated calls under PECR Regulation 19, so direct marketing by automated call needs the recipient's specific prior consent. Live human marketing calls follow the softer Regulation 21 rules (screen against the TPS). | Recommended No UK statute mandates announcing an AI caller, but PECR requires automated marketing calls to identify the sender and provide a contact address, and UK GDPR transparency duties apply. | Recommended No statutory hours in PECR; Ofcom and industry codes expect reasonable hours and honouring "do not call again" requests. | Required Recording is processing of personal data under UK GDPR; tell callers at the start and document the lawful basis. Financial firms have additional FCA recording duties. | 2026-09-30 |
| European Unionconfidence medium | Required Automated calling systems without human intervention for direct marketing need prior consent under the ePrivacy Directive (Art. 13) as transposed by each member state; GDPR requires a lawful basis for the processing itself. | Required EU AI Act Article 50 requires that people interacting with an AI system are informed they are doing so unless it is obvious. Transparency obligations apply from 2 August 2026. Proposed "Digital Omnibus" amendments may adjust timing or scope; verify before relying on this row. | Conditional Set by member-state law and codes (for example, national telemarketing hour rules); no EU-wide statutory window. | Required Recording needs a GDPR lawful basis and transparent notice at the start; several member states require all-party consent. | 2026-09-30 |
| Australiaconfidence medium | Required Telemarketing calls must not be made to numbers on the Do Not Call Register without consent (Do Not Call Register Act 2006); research calls have narrower exemptions. | Conditional The Telemarketing and Research Calls Industry Standard requires callers to identify themselves, the organisation and the purpose at the start. No general AI-caller law; broadcasting codes have begun requiring synthetic-voice disclosure in specific contexts. | Required Telemarketing calls only Monday to Friday 9 a.m. to 8 p.m. and Saturday 9 a.m. to 5 p.m. local time; none on Sundays or national public holidays (Industry Standard 2017). | Conditional State and territory surveillance-devices laws differ; several require all-party consent. Announce recording at the start. | 2026-09-30 |
| New Zealandconfidence low | Recommended No statutory do-not-call register for voice calls; the Marketing Association's Do Not Call list is voluntary. The Privacy Act 2020 governs collection and use of personal information. | Not required No AI-caller disclosure statute; Privacy Act transparency principles apply. | Recommended Industry code expectations only. | Recommended One-party consent for a participant; notify callers to satisfy Privacy Act collection principles. | 2026-09-30 |
Frequently asked
Does drive-thru voice AI actually work at scale?
Public press reports deployments across hundreds of locations at several large quick-service chains in the United States, alongside a well-reported pause in 2025 after customer complaints and prank orders, and a high-profile trial that ended. Treat the numbers as reported, and run your own lane test with real cars before you decide.
What makes drive-thru audio harder than phone audio?
The microphone is outdoors, the customer is in a car with an engine, wind, a radio and passengers, and the agent's own voice comes back through the speaker post. The agent needs barge-in gated on real customer speech, resistance to fabricated words on noise, and a fast hand-off when confidence drops.
Will it push upsells on every customer?
It should offer what you configure, once, and remember a decline within the order. Reported deployments cite check-size gains from consistent suggestion; whether that is good for your brand is your call, and the rule set should be yours.
Can the crew still take over?
Yes, and the hand-off is the most important behaviour to test. A good hand-off puts a crew voice on the speaker within a few seconds with the partial order on their screen; a bad one makes the customer repeat the order.
Related
- Restaurants #1
- Restaurants #2
- Restaurants #4
- Restaurants #5
- Best practice
- Best practice
- Best practice
- Best practice
- Anti-pattern
- Anti-pattern
- Anti-pattern
- Anti-pattern
- Market
- Market