Bluejay / Labs
Open navigation

Working paper / 2026

MIVAS

Multi-Industry Voice Agent Simulation Bench

Abstract

Multi-Industry Voice Agent Simulation Bench (MIVAS) is an indicator of voice AI performance across economic sectors. MIVAS is comprised of tasks and verifiers with granular rewards executed in production-grade multi-agent RL environments across industries where voice AI is being adopted the fastest. Each task places a model harness inside a multi-agent graph with production-length prompts, tools, and isolated state. Deterministic verifiers score tool use, handoffs, and final database state; pass@1 and passk measure single-run capability and repeated-run reliability.

GitHub repository ↗

Clone the repository, pair a model harness with an industry task suite, and reproduce the evaluation environment.

Legal / C1-M2

C1-M2: Identification fails twice

mediumReception and routing· Curtis Beaumont

What this task measures

The caller is on the line with Halverson & Reed for one locked request: Identification fails twice. They keep repeating that ask through a callback offer, a message left for someone else, or advice on a different matter. We included it as a medium reception and routing case because intake has to wait until the caller is classified; messages, wrong numbers, and represented parties are everyday reception work. The agent succeeds only if it stays on that request, reaches the outcome by staying on reception, and leaves the record untouched: no appointment, ticket, or note the caller never asked for.

Performance

One model harness at a time, with a representative pass and fail conversation when both exist. OpenAI Realtime 2.1 is the default view.

Recording

Conversation cost $0.067

0:00 / 1:08 · audio · trace 803836 · click a span to play, drag to scrub, scroll to pan

Audio

Speech Audio

Agent Audio

Tool Calls

Errors

0:000:100:200:300:400:501:00
0:00

Transcript

How a conversation passes

Combined pass is the logical and of three deterministic gates. Tools, handoff path, and final database state must all pass on the same conversation. One miss fails the call.

This sample fails because final database state did not match. Combined cannot pass unless every gate does.

1. Tools

Expected versus actual tool calls on this conversation.

Expected

Actual

  • escalate_to_humanescalate_to_humanreason_code=caller_requesthit

    Parameters

     
    + reason_code: caller_request

    Output

     
    + ok: true

Open a row for parameters and outputs. Required tools must appear. Extra industry tools are allowed.

2. Handoff path

Full industry architecture, same graph as the industry page. The left diagram marks the expected specialist path in green. The right diagram marks the path this conversation actually took in green, because it matched.

Expected path

stay on reception

Actual path

stay on reception

3. Final database state

Rows that differ between the expected hangup and the actual hangup.

escalations · 1

Expected

Actual

- reason_code: identity_failed
+ reason_code: caller_request

Scores across model harnesses

Model harnesspass@1pass5Runs

Task definition

Locked caller script. This is the technical definition. Copy it as written.

Open with exactly: "I got hurt at work and I want to talk to a lawyer." Every time you are asked for your name and phone number, say exactly: "Curtis Beaumont, five five five, one two." Never give ten digits. Never invent the missing digits. If asked about what happened, who this is against, which state, or when it happened, say exactly: "I'd rather talk to a person about that." If they offer only a message or a callback instead of a person, say exactly: "No, I need a person. I don't have a full number." Stay until they confirm a person will take the call because they could not get a usable callback number. A spoken 'we'll have someone call you' without sending you to a person is not enough. When told a person will take it from here, say exactly: "Okay, thank you." and end the call.

Caller traits

full_name
Curtis Beaumont
phone
555-555-0012