Working paper / 2026
Multi-Industry Voice Agent Simulation Bench
Abstract
Multi-Industry Voice Agent Simulation Bench (MIVAS) is an indicator of voice AI performance across economic sectors. MIVAS is comprised of tasks and verifiers with granular rewards executed in production-grade multi-agent RL environments across industries where voice AI is being adopted the fastest. Each task places a model harness inside a multi-agent graph with production-length prompts, tools, and isolated state. Deterministic verifiers score tool use, handoffs, and final database state; pass@1 and passk measure single-run capability and repeated-run reliability.
Clone the repository, pair a model harness with an industry task suite, and reproduce the evaluation environment.
Legal / C5-E2
What this task measures
The caller is on the line with Halverson & Reed for one locked request: Declines after fee disclosure. They keep repeating that ask through a callback offer, a message left for someone else, or advice on a different matter. We included it as an easy fees and booking case because fee disclosure and a two-step evaluation booking are the commercial close of intake, and they have to be spoken from the firm’s own figures. The agent succeeds only if it stays on that request, reaches the outcome by following the conflict screening then intake then scheduling path, and leaves the record untouched: no appointment, ticket, or note the caller never asked for.
One model harness at a time, with a representative pass and fail conversation when both exist. OpenAI Realtime 2.1 is the default view.
Recording
Conversation cost $0.350
0:00 / 3:27 · audio · trace 804046 · click a span to play, drag to scrub, scroll to pan
Audio
Speech Audio
Agent Audio
Handoffs
Tool Calls
Errors
Transcript
How a conversation passes
Combined pass is the logical and of three deterministic gates. Tools, handoff path, and final database state must all pass on the same conversation. One miss fails the call.
Combined
Fail
This sample fails because tools and final database state did not match. Combined cannot pass unless every gate does.
Expected versus actual tool calls on this conversation.
Expected
Actual
Parameters
- practice_area: auto_accident
Output
None recorded.
None recorded.
Parameters
+ full_name: Percival Ndiaye
+ phone: 4045550192
Output
+ ok: true
Parameters
+ from: reception
+ to: screening
Output
None recorded.
None recorded.
Parameters
+ opposing_party: Summit Hauling
Output
+ ok: true
Parameters
+ practice_area: auto_accident
Output
+ ok: true
Parameters
+ practice_area: auto_accident
+ state: CA
Output
+ ok: true
Parameters
+ incident_date: 2026-07-10
+ practice_area: auto_accident
+ state: CA
Output
+ ok: true
Parameters
+ from: screening
+ to: intake
Output
None recorded.
None recorded.
Parameters
+ incident_date: 2026-07-10
+ practice_area: auto_accident
+ state: CA
+ summary: Caller reports being rear-ended by a Summit hauling truck and is considering speaking with a lawyer. Additional details about location and injuries/losses pending.
Output
+ ok: true
Parameters
+ from: intake
+ to: scheduling
Output
None recorded.
None recorded.
Parameters
+ earliest_date: 2026-08-23
+ practice_area: auto_accident
+ state: CA
Output
+ ok: true
Open a row for parameters and outputs. Required tools must appear. Extra industry tools are allowed. Missing: hold_evaluation. Extra: lookup_caller, transfer_to_screening, check_conflict, check_practice_area, check_jurisdiction, calculate_filing_deadline, transfer_to_intake, record_intake, transfer_to_scheduling, find_evaluation_slots.
Full industry architecture, same graph as the industry page. The left diagram marks the expected specialist path in green. The right diagram marks the path this conversation actually took in green, because it matched.
Expected path
transfer_to_screening → transfer_to_intake → transfer_to_scheduling
Actual path
transfer_to_screening → transfer_to_intake → transfer_to_scheduling
Rows that differ between the expected hangup and the actual hangup.
missing · holds · HR-EVAL-3092
Expected
Actual
- caller_id: c_new
- kind: evaluation
- practice_area: auto_accident
- slot_id: s_110
- summary: Tom Eckhart on 2026-08-11 at 09:00. The evaluation is free. If the firm takes the case there is no fee unless it wins: 33.33% of any recovery before a lawsuit is filed, 40% if a lawsuit is filed.
- token: HR-EVAL-3092
intakes · 1
Expected
Actual
- summary: Rear-end collision.
+ summary: Caller reports being rear-ended by a Summit hauling truck and is considering speaking with a lawyer. Additional details about location and injuries/losses pending.
Locked caller script. This is the technical definition. Copy it as written.
Open with exactly: "I was rear-ended by a Summit Hauling truck and I'm thinking about talking to a lawyer." When asked when, say exactly: "It was July tenth, twenty twenty six." If asked to book a free evaluation before any fee numbers, say yes and look at times. When times are offered, take the first slot so they can hold it. After they have held a slot and read back that fee, say exactly: "That's more than I expected — I'll call back later." Do not confirm booking. Do not ask for a staff member. End after declining to book.