Working paper / 2026
Multi-Industry Voice Agent Simulation Bench
Abstract
Multi-Industry Voice Agent Simulation Bench (MIVAS) is an indicator of voice AI performance across economic sectors. MIVAS is comprised of tasks and verifiers with granular rewards executed in production-grade multi-agent RL environments across industries where voice AI is being adopted the fastest. Each task places a model harness inside a multi-agent graph with production-length prompts, tools, and isolated state. Deterministic verifiers score tool use, handoffs, and final database state; pass@1 and passk measure single-run capability and repeated-run reliability.
Clone the repository, pair a model harness with an industry task suite, and reproduce the evaluation environment.
Customer Support / T4-M4
What this task measures
The caller is on the line with Kestrel Electronics for one locked request: Quote a Plus to Total upgrade. They keep repeating that ask through a callback offer, a different ticket, or an unrequested fix. We included it as a medium membership case because status, a prorated upgrade, and a cancellation that offers at most one save are the membership desk in miniature. The agent succeeds only if it stays on that request, reaches the outcome by following the verification then membership path, and leaves the record untouched: no appointment, ticket, or note the caller never asked for.
One model harness at a time, with a representative pass and fail conversation when both exist. OpenAI Realtime 2.1 is the default view.
Recording
Conversation cost $0.313
0:00 / 3:18 · audio · trace 813399 · click a span to play, drag to scrub, scroll to pan
Audio
Speech Audio
Agent Audio
Handoffs
Tool Calls
Errors
Transcript
How a conversation passes
Combined pass is the logical and of three deterministic gates. Tools, handoff path, and final database state must all pass on the same conversation. One miss fails the call.
Combined
Fail
This sample fails because final database state did not match. Combined cannot pass unless every gate does.
Expected versus actual tool calls on this conversation.
Expected
Actual
Parameters
+ fee: how much is Total and upgrading from Plus
Output
+ ok: true
Parameters
+ from: reception
+ to: verification
Output
None recorded.
None recorded.
Parameters
+ full_name: Selena Cortez
+ phone: 5415550171
Output
+ ok: true
Parameters
card_last4: 1156
card_last4: 1156
postal_code: 97035
postal_code: 97035
Output
+ ok: true
Parameters
None recorded.
None recorded.
Output
+ ok: true
Parameters
+ from: verification
+ to: membership
Output
None recorded.
None recorded.
Parameters
None recorded.
None recorded.
Output
+ ok: true
Parameters
+ reason: caller done
Output
None recorded.
None recorded.
Open a row for parameters and outputs. Required tools must appear. Extra industry tools are allowed. Extra: get_fee, end_call.
Full industry architecture, same graph as the industry page. The left diagram marks the expected specialist path in green. The right diagram marks the path this conversation actually took in green, because it matched.
Expected path
transfer_to_verification → transfer_to_membership
Actual path
transfer_to_verification → transfer_to_membership
Rows that differ between the expected hangup and the actual hangup.
extra · holds · KE-UPG-5512
Expected
Actual
+ customer_id: cust_selina
+ kind: upgrade
+ summary: Upgrading to Kestrel Total is $85.00 today, the difference between Plus and Total, prorated over the 6 months left on the current year. It charges to the card ending 1156.
+ token: KE-UPG-5512
Locked caller script. This is the technical definition. Copy it as written.
Open with exactly: "What would it cost to upgrade my Plus membership to Total today?" Give identity when asked. You only want the prorated upgrade quote. Do not confirm the charge. Hang up once they say the amount due. Once they have finished what you asked and confirmed it, say exactly: "Thank you, that's all I needed." and end the call. Do not start a second request.