Working paper / 2026
Multi-Industry Voice Agent Simulation Bench
Abstract
Multi-Industry Voice Agent Simulation Bench (MIVAS) is an indicator of voice AI performance across economic sectors. MIVAS is comprised of tasks and verifiers with granular rewards executed in production-grade multi-agent RL environments across industries where voice AI is being adopted the fastest. Each task places a model harness inside a multi-agent graph with production-length prompts, tools, and isolated state. Deterministic verifiers score tool use, handoffs, and final database state; pass@1 and passk measure single-run capability and repeated-run reliability.
Clone the repository, pair a model harness with an industry task suite, and reproduce the evaluation environment.
Legal / C2-M2
What this task measures
The caller is on the line with Halverson & Reed for one locked request: Refuses to name the other side. They keep repeating that ask through a callback offer, a message left for someone else, or advice on a different matter. We included it as a medium conflicts and barred callers case because conflict screening before facts is the firm’s hard rule, including keeping existing-client names off the line. The agent succeeds only if it stays on that request, reaches the outcome by taking the conflict screening path, and leaves the record untouched: no appointment, ticket, or note the caller never asked for.
One model harness at a time, with a representative pass and fail conversation when both exist. OpenAI Realtime 2.1 is the default view.
Recording
Conversation cost $0.104
0:00 / 1:48 · audio · trace 803896 · click a span to play, drag to scrub, scroll to pan
Audio
Speech Audio
Agent Audio
Handoffs
Tool Calls
Errors
Transcript
How a conversation passes
Combined pass is the logical and of three deterministic gates. Tools, handoff path, and final database state must all pass on the same conversation. One miss fails the call.
Combined
Pass
All three gates matched, so the conversation counts as a pass.
Expected versus actual tool calls on this conversation.
Expected
Actual
Parameters
full_name: Elton Bragg
full_name: Elton Bragg
+ phone: 7045550255
Output
+ ok: true
Parameters
+ from: reception
+ to: screening
Output
None recorded.
None recorded.
Parameters
+ reason_code: conflict_review
Output
+ ok: true
Open a row for parameters and outputs. Required tools must appear. Extra industry tools are allowed.
Full industry architecture, same graph as the industry page. The left diagram marks the expected specialist path in green. The right diagram marks the path this conversation actually took in green, because it matched.
Expected path
transfer_to_screening
Actual path
transfer_to_screening
The actual hangup matches the expected hangup.
Locked caller script. This is the technical definition. Copy it as written.
Open with exactly: "I was hurt on someone's property and I want to know my options." Every time you are asked who this would be against, say exactly: "I'm not comfortable saying that until I know you'll take it." Never name anyone. Never name a state or a date, and never describe what happened. If they start taking a written account or offer an appointment, say exactly: "I'm not comfortable going further until I know you'll take it." Stay until they confirm a person will review this. A spoken promise that they will look into it, without transferring you to a person, is not enough. When told a person will review this, say exactly: "Okay, thank you." and end.