Working paper / 2026
Multi-Industry Voice Agent Simulation Bench
Abstract
Multi-Industry Voice Agent Simulation Bench (MIVAS) is an indicator of voice AI performance across economic sectors. MIVAS is comprised of tasks and verifiers with granular rewards executed in production-grade multi-agent RL environments across industries where voice AI is being adopted the fastest. Each task places a model harness inside a multi-agent graph with production-length prompts, tools, and isolated state. Deterministic verifiers score tool use, handoffs, and final database state; pass@1 and passk measure single-run capability and repeated-run reliability.
Clone the repository, pair a model harness with an industry task suite, and reproduce the evaluation environment.
Legal / R-H2
What this task measures
The caller is on the line with Halverson & Reed for one locked request: Asks to delete the recording. They keep repeating that ask through a callback offer, a message left for someone else, or advice on a different matter. We included it as a hard clients and refusals case because existing-client status and clean refusals are how the firm stays inside its own matters, off the merits. The agent succeeds only if it stays on that request, reaches the outcome by staying on reception, and leaves the record exactly as this case requires: the right row written or cleared, and nothing extra.
One model harness at a time, with a representative pass and fail conversation when both exist. OpenAI Realtime 2.1 is the default view.
Recording
Conversation cost $0.101
0:00 / 1:28 · audio · trace 804118 · click a span to play, drag to scrub, scroll to pan
Audio
Speech Audio
Agent Audio
Tool Calls
Errors
Transcript
How a conversation passes
Combined pass is the logical and of three deterministic gates. Tools, handoff path, and final database state must all pass on the same conversation. One miss fails the call.
Combined
Fail
This sample fails because tools and final database state did not match. Combined cannot pass unless every gate does.
Expected versus actual tool calls on this conversation.
Expected
Actual
Parameters
full_name: Vince Holloway
full_name: Vince Holloway
+ phone: 7025550133
Output
+ ok: true
Parameters
- for_whom: Daniel Okonkwo
Output
None recorded.
None recorded.
Parameters
+ reason_code: caller_request
Output
+ ok: true
Open a row for parameters and outputs. Required tools must appear. Extra industry tools are allowed. Missing: take_message.
Full industry architecture, same graph as the industry page. The left diagram marks the expected specialist path in green. The right diagram marks the path this conversation actually took in green, because it matched.
Expected path
stay on reception
Actual path
stay on reception
Rows that differ between the expected hangup and the actual hangup.
missing · messages · 1
Expected
Actual
- caller_id: c_new
- for_whom: Daniel Okonkwo
- id: 1
- message: Callback requested.
Locked caller script. This is the technical definition. Copy it as written.
Open with exactly: "Delete this recording right now." You have two jobs after they say they cannot control the recording: they look you up, and they take a written message for Daniel Okonkwo that the packet arrived and you need a callback. Stay until they confirm that message is on the file. A spoken 'we'll take care of it' without the write is not enough. Do not accept a person, a new matter, or an appointment until that message is confirmed on the file. When asked what the message is, say exactly: "It's for Daniel Okonkwo. Tell him the packet arrived and I need a callback." After they confirm the message is on the file, say exactly: "And I still want that recording deleted. Put it on the record. I want a person." Then accept transfer. When told a person will take it from here after the message is on the file, say exactly: "Okay, thank you." and end.