Bluejay / Labs
Open navigation

Working paper / 2026

MIVAS

Multi-Industry Voice Agent Simulation Bench

Abstract

Multi-Industry Voice Agent Simulation Bench (MIVAS) is an indicator of voice AI performance across economic sectors. MIVAS is comprised of tasks and verifiers with granular rewards executed in production-grade multi-agent RL environments across industries where voice AI is being adopted the fastest. Each task places a model harness inside a multi-agent graph with production-length prompts, tools, and isolated state. Deterministic verifiers score tool use, handoffs, and final database state; pass@1 and passk measure single-run capability and repeated-run reliability.

GitHub repository ↗

Clone the repository, pair a model harness with an industry task suite, and reproduce the evaluation environment.

Methodology

MIVAS, the Multi-Industry Voice Agent Simulation Bench, is a protocol for scoring speech-to-speech models inside production-shaped voice systems. The sections below are one argument: how we built a world, how we placed a caller in it, how we hosted the model, and how we decided a conversation passed.

00

Overview

MIVAS task execution
MIVAS Bench evaluation process: a Digital Human speaks with a multi-agent voice architecture, then deterministic verifiers score tools, handoffs, and database state conjunctively.

MIVAS scores speech-to-speech models inside production-shaped voice systems. We designed three industry worlds as specialist graphs, each with production-length prompts, tools, and isolated seed state. A voice agent harness hosts that graph for a given provider. Bluejay Digital Humans then place locked callers on the line. Every case states who is calling, what they need, which tools must fire, which specialists must take the call, and what the database must look like at hangup. Each case runs five independent times, each with its own conversation and its own SQLite file.

After hangup, three deterministic verifiers score the same conversation. Tool adherence checks that the required calls fired with the locked arguments. Handoff adherence checks that the specialist path appeared in order. Database-state adherence checks the write-bearing tables against expected state. Combined pass is the logical and of all three: a correct write cannot hide a wrong route, and a correct hop cannot hide a missing tool. Pass@1 is whether one conversation clears that bar. Pass5 asks whether the same capability holds across five trials. Native speech-to-speech systems are the systems under test. A cascaded STT, LLM, and TTS stack is the baseline.

01

Why this protocol

We started from a gap. Existing voice benches already measure live audio, tool use, and sometimes a final database. None of them treat multi-agent routing, provider-native handoffs, and conjunctive verification as one job.

BenchmarkMulti-agent topologyHandoff verificationConjunctive verificationLive adaptive voiceNative S2SMulti-industryStateful toolsFinal-state verifierTool-adherence verificationRepeated-run reliability
MIVAS BenchYesYesYesYesYesYesYesYesYesYes
EVANoNoNoYesYesYesYesYesPartialYes
τ-VoiceNoNoNoYesYesYesYesYesNoPartial
VAmoS BenchNoNoNoYesYesNoYesNoPartialYes
Full-Duplex-Bench v3NoNoNoNoYesPartialPartialNoYesNo
VoiceAgentBenchNoNoNoNoNoPartialNoNoPartialNo

Yes · Partial · No. Source: mivas-bench README benchmark comparison.

Voice AI is already answering phones in dermatology groups, plaintiff firms, and national retail support centers. Those deployments are specialist graphs. Handoffs are invisible to the caller. Authorization, policy, and fee math are the work, not decoration.

A model can book the right appointment on the wrong specialist, or take the right route and never fire the write the office needs. Collapsing the call into one judge score hides which of those happened. MIVAS is the only row that stays yes across topology, handoff, and conjunctive checks.

02

Harness and pack

We split the repository so the model runtime and the industry world can be paired without rewriting either side. A scored environment is always one harness plus one pack.

Harness plus industry pack

The harness is the provider adapter. It carries bidirectional audio, builds the agents declared by the blueprint, performs native handoffs and session operations, and POSTs industry tools. It does not own policy, prices, or schema.

The pack owns the multi-agent blueprint, production-length prompts, tool definitions, the database schema and seed, the FastAPI state service, and the task suite. agent_blueprint.json is the interface. The same OpenAI Realtime harness can run Straus, Halverson & Reed, or Kestrel. The same Kestrel pack can run against Gemini Live or Grok.

control-industry is the setup fixture, not a scored pack. A new harness must receive a call, construct the declared agents, complete a native handoff, invoke a tool, and persist an appointment before it is eligible for healthcare, legal, or customer support. Control results never enter the leaderboard.

03

Industry environments

We wrote three complete companies from production voice lines we already knew, then replaced every name, chart, and order so the world is fictional, versioned, and shared across harnesses. Healthcare, legal, and customer support are the sectors where those lines are already live.

Each pack is a fully specified production environment: a specialist graph, node-level prompts, domain tools, seeded state, and locked callers. The rest of the protocol holds that world constant and varies the model runtime. The one-pagers below are the worlds. Full fixtures and node prompts live on the industries pages.

Healthcare: Straus Dermatology

Healthcare handoff graph

We built the healthcare pack from years of evaluating voice front desks at dermatology groups, then wrote Straus Dermatology as a complete hypothetical practice and Robin as its inbound line. Straus is a 160-location, 380-provider group across nine states. The name rhymes with house. Robin is one continuous person from hello to goodbye: warm, Northeast-neutral. Fixture offices are Park Avenue (cosmetic), Brooklyn Heights (cosmetic), and Windermere (medical). The clock is pinned to 19 August 2026 at noon. Open slots repeat at 9:00, 11:30, and 14:00.

The architecture is the one we keep seeing on the strongest deployments: a seven-node DAG. Reception greets and answers public office facts through list_locations, then routes. Identity is the PHI gate: name and date of birth, then verify_identity and get_patient_summary before any protected desk. Scheduling, coverage, cosmetic, billing, and clinical each own a desk. Handoffs are invisible. Escalation is one global tool, transfer_to_human. Specialists stay downstream of reception. Scheduling and cosmetic are sinks. Billing and clinical may hop forward to scheduling and never back to identity.

The measurement surface is the work that stalls a front desk: new-patient access, appointment management, office-level plan acceptance, approved-table cosmetic quotes, billing resolution, and clinical liaison that never reads results or approves a refill. Tools match those flows. classify_visit_request, check_plan_accepted, and book_appointment carry a new patient onto the board. quote_cosmetic_service holds the approved table for Botox, filler, peel, and microneedling. explain_charge and request_fee_waiver sit on billing. Refill hard stops are isotretinoin, controlled substances, and biologics. A typical new-patient path is classify, coverage, slot bind, book, and a confirmation text: several writes inside a front-desk latency budget.

Legal: Halverson & Reed

Legal handoff graph

We built the legal pack from intake work with plaintiff firms and from the ethics rules those firms have to keep on a live line. Halverson & Reed is the replica: a contingency practice taking injury, employment, and consumer matters across nine states. Hal is the inbound front desk. Talking to Hal leaves the caller a prospective client until an attorney says otherwise. An attorney makes every decision to take or decline a matter.

Hal is a five-node chain, simpler than healthcare's fan-out. Reception classifies the caller. Screening runs conflict, practice area, jurisdiction, and filing deadline, in that order. Intake writes the matter and offers the packet. Scheduling books the free evaluation through a hold-and-confirm gate. Client services reports status on the firm's own matters and never the merits. The graph does not fan out. The difficulty lives in the work each of those tasks requires.

Conflict screening runs before any facts of the matter, following ABA Rule 1.18. Practice area and jurisdiction are two gates, not one. The default footprint is nine states. Medical malpractice is licensed only in Florida, Georgia, and New York. Workers' compensation is licensed only in California, Florida, Georgia, and Texas. calculate_filing_deadline returns expired, urgent, or ok, and the agent reads the date as returned. Hal escalates with a reason code such as conflict, represented_party, adverse_party, practice_area, jurisdiction, or deadline_review. Fee figures come only from hold_evaluation. Contingency is thirty-three and a third percent before filing and forty percent after. Workers' compensation is twenty percent. Consumer is hourly at $175. Two-step writes use fixed tokens (HR-EVAL-3092, HR-CANC-7715) so read-back is checkable. A full new-matter path crosses lookup_caller, the four screening checks, record_intake, send_intake_packet, find_evaluation_slots, hold_evaluation, and confirm_evaluation.

Customer support: Kestrel Electronics

Customer support handoff graph

The customer-support pack comes from retail voice lines we have seen in production, where the work is order status, returns math, delivery changes, membership, and impersonation reports. Kestrel Electronics is the replica: a national consumer-electronics retailer of about a thousand stores, founded in Wexley, Ohio in 1971, support center in Oregon. Every call opens with an AI disclosure and a recorded-line disclosure. Callers still use acquired names (Sound Harbor, Bellwether Mobile, Aurelian Audio, Coastline Kitchen & Home, Sagebrush Outdoor). The agent treats those as Kestrel.

The inbound line is one conversation, served by seven instruction sets. Reception answers public facts. Verification is the order-bound identity gate: name plus phone or order number, then the order ZIP and card last four, in one question. Orders, returns, service, and membership sit behind it. Fraud is a separate desk for impersonation calls, and it is ungated on purpose. Demanding a ZIP and a card from a frightened impersonation caller is the scammer's own move. TechCrew is the service arm. Kestrel Plus is $29.99 a year. Kestrel Total is $199.99 a year and adds TechCrew Protect on most purchases for up to two years.

Policy math is pinned in the state API so the spoken figure can be checked. Return windows are 15 days standard and 60 days for members, with activatable devices at 14 days on every tier. Restocking is $45 on opened phones and similar devices, and 15 percent on drones and special-order products, with a $0 fee in eight states including Ohio. Delivery changes stop at a 48-hour boundary. Coverage is a four-rung ladder: TechCrew Protect, Kestrel Total, manufacturer warranty, then a $39.99 bench diagnostic. A typical write path is verify_identity, a quote, a spoken read-back, and a confirm token (KE-RTN-4417, KE-DLV-3390, KE-CXL-7708). Glen Aldridge is the headline trap: a Total member returning a Solstice X5, sold sixty-day returns, with a fourteen-day activatable window.

Prompts are production-length operating specifications, not shortened benchmark instructions. Shared CORE rules appear in every node: identity, invisible handoffs, spoken commitments, and refusal scripts. Specialists never re-greet. The only transfer announced out loud is the human escalation. Solid edges are specialist transfers. Dashed edges escalate to a human.

04

Isolated state

Once the company existed, we still needed every conversation to start from the same seed and end with a dump that only that conversation could have written. We keyed the database to the Bluejay simulation result id.

Per-call SQLite isolation

The Digital Human places that id on the CHIRP upgrade as X-Simulation-Result-Id, or on the SIP INVITE for LiveKit workers. The harness stamps every industry tool POST with X-Mivas-Call-Id. Missing the header is a 400, except for the local shared-DB fallback used by --check.

First touch copies schema.sql and seed.sql into calls/{id}.db. Later tools reopen that file. Harnesses are dumb pipes: they POST {TOOL_SERVER_URL}/tools/{name} and return the envelope from tools.json. Handoff and session tools never hit the state API.

On Kubernetes, evals do not GET the public hostname for /state. An ALB would pick a random replica. At hangup the replica that owned the socket writes the snapshot to S3. A call that never invoked a tool still has a defined initial state: first GET /state for that id lazy-creates the file from seed.

05

Digital humans

A world without a specified caller is a demo. We wrote Bluejay Digital Humans as locked users: an identity, one objective, traits, and scripted replies that keep the measurement on the task.

They are live, adaptive voice callers, not IVR script-readers. They ask again when the office-level answer is missing. They decline a waitlist, a person, or a confirmation text when the case forbids it. Creativity is pinned so avoidable variance does not wash out the score.

Healthcare C1-E1 is the teaching case. Dana Whitfield is a new patient who wants a yes or no for Aetna at Park Avenue, not for the practice. If the agent says it takes Aetna without naming the office, she asks which office was checked. The required tool is check_plan_accepted with carrier=aetna and location_id=loc_park_ave. The expected hop is coverage. The expected database is the unchanged seed. This is a read-only case, which is why tool adherence cannot be folded into final state.

Twelve cases in each pack replay a selected script under background noise or a degraded signal. The intent, tools, and expected state stay the same. The channel does not.

06

Task suites

We locked 72 cases in each industry: 60 base tasks and 12 audio variants, balanced across easy, medium, and hard, then reviewed and replayed every case against fresh seed state.

Difficulty, per industry

24 easy, 24 medium, 24 hard. The same split is used in healthcare, legal, and customer support.

Easy24/72
Medium24/72
Hard24/72

Audio condition, per industry

60 clean-line cases, then 12 locked variants that replay selected cases under cafe or office noise, or a degraded signal.

Perfect60/72
Background noise6/72
Bad signal6/72

Five categories cover the principal workflows. A sixth R category tests regulation, refusal, escalation, impersonation, and adversarial asks. The R label is industry-specific: regulatory adherence in healthcare and customer support, clients and refusals in legal.

Every task.json states the caller, the objective, the traits, the scripted replies, the expected handoff path, the required tools with constrained arguments, and the exact final database. Identifiers encode the suite. C1-E1 is category 1, easy, first case. C1-E1-BG and C1-E1-SIG are its noise and signal variants. Customer support uses T1 to T5 the same way.

The catalog at Tasks is the public lockfile. Copy a single case or the visible set as a replication prompt.

07

Voice agent harnesses

The same locked cases then run on twelve provider runtimes. We treat the harness as part of the system under test, because a production voice agent is the model plus the runtime that can actually host the graph.

Three handoff implementations
RuntimeStackTransportHandoff
GPT Live 1 (astra-medium)

gpt-live-1 + gpt-6-astra

Speech-to-speechCHIRP WebSocketNative RealtimeAgent graph
GPT Live 1 (sol-low)

gpt-live-1 + gpt-5.6-sol

Speech-to-speechCHIRP WebSocketNative RealtimeAgent graph
Gemini 3.8 Live Extended Thinking

gemini-3.8-live-extended-thinking

Speech-to-speechSIPNew Gemini Live socket
Gemini 3.8 Live

gemini-3.8-live

Speech-to-speechSIPNew Gemini Live socket
OpenAI Realtime 2.1

gpt-realtime-2.1

Speech-to-speechCHIRP WebSocketNative RealtimeAgent graph
Grok Voice 2

grok-voice

Speech-to-speechCHIRP WebSocketSoft session.update
Qwen Audio Realtime

qwen-audio-3.0-realtime-plus

Speech-to-speechCHIRP WebSocketSoft session.update
Amazon Nova Sonic 2

amazon.nova-sonic-v2

Speech-to-speechCHIRP WebSocketNew Bedrock stream
Gemini Flash Live 3.1

gemini-3.1-flash-live

Speech-to-speechSIPNew Gemini Live socket
Cascaded Baseline

flux-general-en + gpt-4.1 + eleven_flash_v2_5

Cascaded STT / LLM / TTSSIPIn-framework agent switch
OpenAI Realtime 2.1 Mini

gpt-realtime-2.1-mini

Speech-to-speechCHIRP WebSocketNative RealtimeAgent graph
Gemini 2.5 Flash Native Audio

gemini-2.5-flash-native-audio

Speech-to-speechSIPNew Gemini Live socket

Providers do not expose the same multi-agent primitive, so the industry DAG is implemented three ways. OpenAI Realtime and Cascaded Baseline use a native agent graph. Grok and Qwen keep one WebSocket and swap instructions with session.update. Nova 2 Sonic and Gemini Live open a new stream or socket per specialist, then seed it so the next voice continues mid-call.

CHIRP-family pods expose an ALB WebSocket. Gemini and Cascaded Baseline register a LiveKit worker; the simulation engine dials SIP. Speak-first, echo, and barge-in stay in the harness. OpenAI semantic VAD only answers caller audio, so a Digital Human that waits would deadlock without a response.create nudge. Those details are why two speech-to-speech models with similar published capability can diverge on the same locked case. See Model Harnesses.

08

Conjunctive verification

A single artifact of the call cannot stand in for the whole job. Final state cannot see a read. A correct tool sequence cannot see a wrong specialist. Combined pass is the logical and of three independent checks on the same conversation.

Combined pass

A conversation counts only when all three deterministic gates pass on the same call. Combined pass is the logical and, not a weighted score.

1. Tools

Required industry calls fire with the constrained arguments the task locks. Extra industry tools are allowed.

2. Handoff

Expected specialist transfers appear in order. Extra hops are allowed once the required path is a subsequence.

3. Final DB

Hangup snapshot of write-bearing tables matches the expected state produced from the same seed.

Combined

Pass

A database-state diff only sees what was written. Many locked cases are satisfied by reads: plan acceptance, a filing deadline, order status, eligibility. If the agent answers from memory or a practice-wide guess, the seed is unchanged and a state-only verifier would pass. That is why we score tool adherence. The required calls must fire with the locked arguments, so a coverage question that never called check_plan_accepted fails even when the hangup snapshot looks clean.

Tools can still fire from the wrong node. If reception books the visit, or coverage writes a field it should never see, the arguments may match and the final tables may match, but the specialist graph has leaked. That is a scoping failure: the runtime let an agent use tools outside its role. Handoff adherence is also how we see a bloated walk. An agent that hops through extra desks on the way from the entry node to the sink has taken an unoptimal path, even if it eventually called the right tools. The gate asks whether the conversation followed the intended route, not merely whether it ended in the right place.

scripts/verify_task_run.py joins a Bluejay result to task.json by case_key or the test_name prefix, then runs the three gates. Tool adherence matches expected calls by name and constrained arguments. Extra industry tools are allowed. Handoff adherence requires the expected hops as an in-order subsequence of transfer tools, sorted by start_offset_ms. An empty expected path always passes, which is how some R refusals stay reception-only. Database-state adherence compares the hangup snapshot to exp_db_state on write-bearing tables. Catalog rows, created_at, and unconstrained prose are ignored.

Combined pass requires all three. A successful write cannot hide a missing read. A correct tool sequence cannot hide a wrong specialist or a wandering route. Transcript quality, latency, and cost stay on the export as diagnostics. Prompt adherence and task completion are auxiliary 1 to 5 judge scores. Premature end is a yes or no judge. None of those enter combined pass. The matching rules, aliases, skip semantics, and worked examples are specified on the conjunctive verifier page.

09

Pass@1 and Pass⁵

A production voice line is judged on whether it can do the same request again. We ran every locked case five times, independently, each with its own result id and its own SQLite file.

How pass@1 and pass5 diverge

Each row is one locked task with five scored trials. pass@1 is c / 5. pass5 is C(c, 5) / C(5, 5): 100% only when all five trials pass. The published pack score is the unweighted mean of those per-task pass5 values.

5/5 passed

pass@1100.0%
pass^5100.0%

4/5 passed

pass@180.0%
pass^50.0%

3/5 passed

pass@160.0%
pass^50.0%

2/5 passed

pass@140.0%
pass^50.0%

1/5 passed

pass@120.0%
pass^50.0%

0/5 passed

pass@10.0%
pass^50.0%

pass@1 is the single-attempt rate: scored conversations that cleared all three gates, among pass and fail only. Miss and void trials drop out of the denominator. For one task this is c / n. For a pack it is that fraction pooled across 360 calls.

pass5 estimates the chance the same agent would succeed on all five attempts of a task. We use the unbiased estimator from tau-bench (Yao et al., 2024): for a task with c passes among n trials, passk = C(c, k) / C(n, k), averaged across tasks. With five trials and k = 5 that is 100% when all five pass and 0% otherwise, so three of five is 60% pass@1 and 0% pass5. That drop is the reliability claim.

HumanEval reports pass@k, the chance that at least one of k samples succeeds: pass@k = 1 - C(n - c, k) / C(n, k). We report passk because a front desk that books the visit once in five calls is not a front desk. Earlier versions of this site used the plug-in estimate (c / n)k, which credits four of five with 32.8%; every published number now uses the unbiased form, and the mivas-bench README states the same definition.

10

Evaluation matrix

The released matrix is the product of that design: twelve completed runtimes, three scored packs, 72 cases, five trials.

Scored runtimes

12

Industry packs

3

Locked tasks

216

Scored conversations

12,960

From locked case to pass@1 and pass^5

Exports keep one row per conversation: task identity, component passes, state differences, transcript, trace, latency, metrics, and estimated cost. This site ingests those CSVs into src/data/ without signed recording URLs. pass@1, pass5, and the three gates can be recomputed from the evidence.

Cross-industry pass@1 is the unweighted mean of the three pack scores. Easy, medium, and hard cells are the unweighted mean of those difficulty scores across packs. Cost per hour is provider list price on scored conversations, from traced token usage or $0.08 per audio minute for Grok. Cascaded Baseline prices Deepgram Flux, GPT-4.1, and ElevenLabs Flash 2.5 together.

11

Reproducibility

A score is only useful if another team can name the revision, the pair, and the evidence. We made each of those first-class.

Companies and records are fictional. Schemas, seeds, prompts, tools, and handoff graphs are versioned. Trials stay separate. Comparisons should identify the repository revision, harness, industry suite, model versions, trial count, concurrency, and any retries. New harnesses start on control-industry. New cases state the fixture, the goal, the required and prohibited conduct, the expected tools, the expected state, and the evidence used to score.

Locally: clone mivas-bench, set HARNESS and INDUSTRY, run uv run python run.py --check, then run.py --apply for the Kubernetes pair.

12

Limitations

The protocol is specific on purpose. That specificity is also the bound. Do not read a MIVAS score as a verdict on every voice system, every accent, or every sector.

Not every implemented harness has completed every industry. Provider telemetry varies. Simulated callers support reproducibility and do not represent every property of human speech. Component scores are evaluation outputs, not an RL training interface. Do not compare scores across task, prompt, verifier, or runtime revisions without qualification.

Judge scores exist because they catch spoken-policy failures the gates cannot see: a correct check_plan_accepted whose script is never spoken, a membership cancel after a second save offer. They inherit the usual limits of LLM-as-judge. This page is the first publication of the protocol. The industry one-pagers, the task lockfile, and the harness adapters are the supporting materials.

Citation

@misc{siddiqi2026mivasbench,
      title={MIVAS Bench: Multi-Industry Voice Agent Simulation Bench},
      author={Faraz Siddiqi and Yash Savalia},
      year={2026},
      url={https://github.com/bluejay-ai-dev/mivas-bench},
}

Shunyu Yao, Noah Shinn, Pedram Razavi, Karthik Narasimhan, “τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains,” 2024, arXiv:2406.12045. Source of the pass^k and pass@k estimators.