Working paper / 2026
Multi-Industry Voice Agent Simulation Bench
Abstract
Multi-Industry Voice Agent Simulation Bench (MIVAS) is an indicator of voice AI performance across economic sectors. MIVAS is comprised of tasks and verifiers with granular rewards executed in production-grade multi-agent RL environments across industries where voice AI is being adopted the fastest. Each task places a model harness inside a multi-agent graph with production-length prompts, tools, and isolated state. Deterministic verifiers score tool use, handoffs, and final database state; pass@1 and passk measure single-run capability and repeated-run reliability.
Clone the repository, pair a model harness with an industry task suite, and reproduce the evaluation environment.
Healthcare / C1-E1-BG
What this task measures
The caller is on the line with Straus Dermatology for one locked request: Aetna coverage at Park Avenue. They keep repeating that ask through a waitlist offer, a callback, or a different errand. We included it as an easy new patient access case because first-contact dermatology work is scored on holding a narrow office question through booking offers, waitlists, and extra yeses. The agent succeeds only if it stays on that request, reaches the outcome by taking the coverage desk path, and leaves the record untouched: no appointment, ticket, or note the caller never asked for.
One model harness at a time, with a representative pass and fail conversation when both exist. OpenAI Realtime 2.1 is the default view.
Recording
Conversation cost $0.689
0:00 / 8:08 · audio · trace 801708 · click a span to play, drag to scrub, scroll to pan
Audio
Speech Audio
Agent Audio
Handoffs
Tool Calls
Errors
Transcript
How a conversation passes
Combined pass is the logical and of three deterministic gates. Tools, handoff path, and final database state must all pass on the same conversation. One miss fails the call.
Combined
Fail
This sample fails because tools did not match. Combined cannot pass unless every gate does.
Expected versus actual tool calls on this conversation.
Expected
Actual
Parameters
+ from: reception
+ to: coverage
Output
None recorded.
None recorded.
Parameters
- carrier: aetna
- location_id: loc_park_ave
Output
- data: {"accepted":true}- ok: true
Parameters
+ location_id: loc_park_ave
+ zip: 10016
Output
+ ok: true
Parameters
+ reason: caller_done
Output
None recorded.
None recorded.
Open a row for parameters and outputs. Required tools must appear. Extra industry tools are allowed. Missing: check_plan_accepted. Extra: list_locations, end_call.
Full industry architecture, same graph as the industry page. The left diagram marks the expected specialist path in green. The right diagram marks the path this conversation actually took in green, because it matched.
Expected path
transfer_to_coverage
Actual path
transfer_to_coverage
The actual hangup matches the expected hangup.
Locked caller script. This is the technical definition. Copy it as written.
You are a new patient calling Straus for the first time. You do not have a provider picked, and you are not booking today. Ask whether they take Aetna at the Park Avenue office specifically — that Manhattan office, not a yes for the whole practice. If they ask which plan, name your carrier: Aetna. If they ask for a provider name, say you do not have one yet and you still need a yes or no for Aetna at Park Avenue itself. If they offer to schedule, check it during booking, handle coverage later, put you on a waitlist, text you, or transfer you to a person, decline and stay on that office question. Do not thank them and end until you hear a clear yes or no for Park Avenue. Do not ask to book, reschedule, or cancel. Do not ask to add insurance, for a callback, a bill, a payment, the patient portal, the street, floor, suite, or financing. A spoken yes for the whole practice is not enough. A spoken yes or no that names the office is enough. Thank them and end once you have that office-level answer. Do not ask them to text you an appointment confirmation.