Bluejay / Labs
Open navigation

Working paper / 2026

MIVAS

Multi-Industry Voice Agent Simulation Bench

Abstract

Multi-Industry Voice Agent Simulation Bench (MIVAS) is an indicator of voice AI performance across economic sectors. MIVAS is comprised of tasks and verifiers with granular rewards executed in production-grade multi-agent RL environments across industries where voice AI is being adopted the fastest. Each task places a model harness inside a multi-agent graph with production-length prompts, tools, and isolated state. Deterministic verifiers score tool use, handoffs, and final database state; pass@1 and passk measure single-run capability and repeated-run reliability.

GitHub repository ↗

Clone the repository, pair a model harness with an industry task suite, and reproduce the evaluation environment.

Healthcare / C4-E1

C4-E1: Botox cost inquiry

easyCosmetic concierge· Bettina Rausch

What this task measures

The caller is on the line with Straus Dermatology for one locked request: Botox cost inquiry. They keep repeating that ask through a waitlist offer, a callback, or a different errand. We included it as an easy cosmetic concierge case because quoted cosmetic policy, deposits, and the cooling-off window have to be spoken before a booking is taken. The agent succeeds only if it stays on that request, reaches the outcome by taking the cosmetic concierge path, and leaves the record untouched: no appointment, ticket, or note the caller never asked for.

Performance

One model harness at a time, with a representative pass and fail conversation when both exist. OpenAI Realtime 2.1 is the default view.

Recording

Conversation cost $0.128

0:00 / 1:48 · audio · trace 801881 · click a span to play, drag to scrub, scroll to pan

Audio

Speech Audio

Agent Audio

Tool Calls

Errors

0:000:100:200:300:400:501:001:101:201:301:40
0:00

Transcript

How a conversation passes

Combined pass is the logical and of three deterministic gates. Tools, handoff path, and final database state must all pass on the same conversation. One miss fails the call.

This sample fails because tools and handoff path did not match. Combined cannot pass unless every gate does.

1. Tools

Expected versus actual tool calls on this conversation.

Expected

Actual

  • transfer_to_cosmeticmissing

    Parameters

    None recorded.

    None recorded.

    Output

    None recorded.

    None recorded.

  • quote_cosmetic_serviceservice=botoxmissing

    Parameters

    - service: botox
     

    Output

    - data: {"range_low":300,"range_high":600}
     
    - ok: true
     
  • end_callreason=caller_doneextra

    Parameters

     
    + reason: caller_done

    Output

    None recorded.

    None recorded.

Open a row for parameters and outputs. Required tools must appear. Extra industry tools are allowed. Missing: transfer_to_cosmetic, quote_cosmetic_service. Extra: end_call.

2. Handoff path

Full industry architecture, same graph as the industry page. The left diagram marks the expected specialist path in green. The right diagram marks the path this conversation actually took in green through the shared prefix, then red from the first wrong hop.

Expected path

transfer_to_cosmetic

Actual path

stay on reception

3. Final database state

The actual hangup matches the expected hangup.

Scores across model harnesses

Model harnesspass@1pass5Runs

Task definition

Locked caller script. This is the technical definition. Copy it as written.

Ask how much Botox costs at the Park Avenue office. You want a price range only. Once you have it, thank them and end the call. Do not book a consult even if they offer one. If they offer a waitlist, financing, a deposit, a text, or a transfer to a person, decline. Do not ask to book, reschedule, or cancel. Do not ask to add insurance, for a callback, a bill, a payment, the patient portal, or the street, floor, or suite. Do not ask them to text you an appointment confirmation.

Caller traits

full_name
Bettina Rausch
preferred_office
Park Avenue