Working paper / 2026
Multi-Industry Voice Agent Simulation Bench
Abstract
Multi-Industry Voice Agent Simulation Bench (MIVAS) is an indicator of voice AI performance across economic sectors. MIVAS is comprised of tasks and verifiers with granular rewards executed in production-grade multi-agent RL environments across industries where voice AI is being adopted the fastest. Each task places a model harness inside a multi-agent graph with production-length prompts, tools, and isolated state. Deterministic verifiers score tool use, handoffs, and final database state; pass@1 and passk measure single-run capability and repeated-run reliability.
Clone the repository, pair a model harness with an industry task suite, and reproduce the evaluation environment.
Healthcare / C3-H1
What this task measures
The caller is on the line with Straus Dermatology for one locked request: Aetna Open Access Elect Choice at Park Avenue. They keep repeating that ask through a waitlist offer, a callback, or a different errand. We included it as a hard coverage and benefits case because carrier answers have to stay tied to one office and one plan, naming that location and that carrier. The agent succeeds only if it stays on that request, reaches the outcome by following the coverage desk then scheduling path, and leaves the record exactly as this case requires: the right row written or cleared, and nothing extra.
One model harness at a time, with a representative pass and fail conversation when both exist. OpenAI Realtime 2.1 is the default view.
Recording
Conversation cost $0.537
0:00 / 6:28 · audio · trace 801842 · click a span to play, drag to scrub, scroll to pan
Audio
Speech Audio
Agent Audio
Handoffs
Tool Calls
Errors
Transcript
How a conversation passes
Combined pass is the logical and of three deterministic gates. Tools, handoff path, and final database state must all pass on the same conversation. One miss fails the call.
Combined
Fail
This sample fails because tools and handoff path and final database state did not match. Combined cannot pass unless every gate does.
Expected versus actual tool calls on this conversation.
Expected
Actual
Parameters
+ from: reception
+ to: coverage
Output
None recorded.
None recorded.
Parameters
- carrier: aetna
- location_id: loc_park_ave
Output
- data: {"accepted":true}- ok: true
Parameters
None recorded.
None recorded.
Output
None recorded.
None recorded.
Parameters
- is_new_patient: true
- urgency: routine
- visit_class: medical
Output
None recorded.
None recorded.
Parameters
location_id: loc_park_ave
location_id: loc_park_ave
+ zip: 10016
Output
+ ok: true
Parameters
- location_ids: ["loc_park_ave"]
Output
None recorded.
None recorded.
Parameters
- appointment_type_code: NP_MED
- end: 2026-08-24T09:30
- location_id: loc_park_ave
- provider_id: prov_chen
- slot_id: slot_loc_park_ave_1
- start: 2026-08-24T09:00
Output
- data: {"status":"booked"}- ok: true
Parameters
- appointment_type_code: NP_MED
- earliest: 2026-08-24
- latest: 2026-09-30
- location_ids: ["loc_park_ave"]
Output
- data: {"status":"added"}- ok: true
Parameters
+ callback_number: +12125550182
+ priority: routine
+ queue: front_desk
Output
+ ok: true
Parameters
+ destination: patient_support_center
+ reason: caller_request
Output
+ ok: true
Open a row for parameters and outputs. Required tools must appear. Extra industry tools are allowed. Missing: check_plan_accepted, transfer_to_scheduling, classify_visit_request, find_slots, book_appointment, join_waitlist. Extra: create_callback_task, transfer_to_human.
Full industry architecture, same graph as the industry page. The left diagram marks the expected specialist path in green. The right diagram marks the path this conversation actually took in green through the shared prefix, then red from the first wrong hop.
Expected path
transfer_to_coverage → transfer_to_scheduling
Actual path
transfer_to_coverage → transfer_to_human
Rows that differ between the expected hangup and the actual hangup.
missing · appointments · 4
Expected
Actual
- appointment_type_code: NP_MED
- description: new patient visit
- end: 2026-08-24T09:30
- id: 4
- location_id: loc_park_ave
- patient_id: empty
- provider_id: prov_chen
- start: 2026-08-24T09:00
- status: booked
missing · waitlist · 1
Expected
Actual
- appointment_type_code: NP_MED
- earliest: 2026-08-24
- id: 1
- latest: 2026-09-30
- patient_id: empty
Locked caller script. This is the technical definition. Copy it as written.
You are a new patient and you are particular about the plan name. When they ask who you are, give your name, date of birth, and mobile once, confirm the readback, and say you are new. Ask whether Straus takes Aetna Open Access Elect Choice at Park Avenue, and say that full plan name, not just the carrier. Once they have answered about Aetna at that office, even if they only confirm the carrier and not the exact plan name, ask to be booked for a rash on your neck. Take the first appointment time they offer. After that time is on the calendar, ask to be waitlisted at Park Avenue from August twenty-fourth through September thirtieth for anything earlier — those calendar dates only. Confirm that first time out loud, then thank them and end the call. Do not ask them to flag the visit. Do not ask to add or change insurance. Do not ask for a callback, to speak to a person, a bill, a payment, the patient portal, a text, financing, or the street, floor, or suite. Stay until they have looked up whether Park Avenue takes Aetna; a guess without a lookup is not enough. A listed opening is not a booking. Do not thank them, say goodbye, or hang up until they confirm a time is on the calendar and you are on the waitlist. If they only list openings, say book the first one. Do not ask them to text you an appointment confirmation.