Bluejay / Labs
Open navigation

Whitepaper / Sep 2026

MIVAS Bench: Building RLVR environments for audio models

How we built a production-shaped, fully deterministic voice benchmark, and why we believe it yields the highest-fidelity training signal available for speech-to-speech models today.

Faraz Siddiqi

Co-founder & CTO, Bluejay

Published

September 2, 2026

MIVAS task execution: a simulated conversation runs through a multi-agent voice architecture and ends in deterministic verification

Abstract

We introduce MIVAS Bench, the Multi-Industry Voice Agent Simulation Bench: a benchmark and RLVR environment suite for speech-to-speech (S2S) models across healthcare, legal, and customer support. MIVAS makes two contributions. The first is multi-agent voice harnesses that reflect how customers build voice systems in production today: each environment is a graph of specialist subagents with mutually exclusive tools, production-length prompts, provider-native handoffs, and isolated per-call state, backed by a fully implemented tool server and database. The second is a conjunctive verifier that makes multi-agent voice tasks deterministically verifiable. It combines a database-state diff, tool-call adherence, and handoff adherence with a logical AND, so read-only tasks are verified without a judge model, and it scores the handoff path against the optimal route, which yields partial rewards for correct-but-inefficient solutions and a full reward only for the optimal one. Every case runs five independent times; pass@1 reports capability and pass5 reports reliability. The bench is open source and open data.

For the abridged version, see my thread on X, which covers the main takeaways. For the architecture overview, see the MIVAS methodology.

01 / Overview

What MIVAS Bench measures

MIVAS Bench is an indicator of voice AI performance across economic sectors. It is a set of tasks and deterministic verifiers, with granular rewards, executed inside production-grade multi-agent environments in the industries where voice AI is being adopted fastest. The initial release covers healthcare, legal, and customer support. For each, we built a voice agent for a fictitious business in extreme detail: a tool server with fully implemented functions, an attached database, and a multi-agent architecture whose tools and scenarios are drawn from real deployments in that industry.

MIVAS makes two contributions to voice benchmarks. The first is realism. Each environment is a graph of specialist subagents, the way enterprises actually build voice agents today, with production-length prompts, provider-native handoffs, and isolated per-call state. The second is verifiable reward. Every conversation is scored by three deterministic checks, database state, tool calls, and handoff path, combined with a logical AND, and the handoff check yields a partial reward that distinguishes an optimal route from a merely successful one. No judge model sits in the scoring path.

The figure at the top of this page is one task executing end to end:

  • Simulated conversation. A Bluejay Digital Human, a locked caller with an identity, one objective, and scripted replies, talks to the agent over live audio.
  • Multi-agent voice architecture. The agent is a directed acyclic graph of specialists. Each node exposes its own toolset; end call, transfer, and escalate are shared session tools available everywhere.
  • Deterministic verification. At hangup, DB-state adherence, handoff adherence, and tool adherence are checked independently and combined with AND. One miss fails the task.

Each locked case runs five independent times. Pass@1 reports single-run capability and pass5 reports whether the same capability holds across all five trials. Native speech-to-speech models are the systems under test; a cascaded STT, LLM, and TTS stack is the baseline. The bench is open source and open data, and every task is packaged as an RLVR environment.

We believe this combination gives MIVAS the highest training signal of any voice benchmark available today. The sections that follow explain why: what existing benchmarks measure, how we designed the environments and tasks, how the verifier works, how the reward is shaped, and what the scores do and do not say.

Can a voice benchmark be production-realistic and fully deterministic at the same time, and can its reward be granular enough to train on?

02 / Current landscape

What existing voice benchmarks measure

Several benchmarks already measure live audio, tool use, and in some cases a final database. Tau Bench established the database-state diff as the deterministic verifier of choice for text agents, and it works well for write-heavy tasks. For read-based tasks (retrieving store hours for a clinic, confirming a plan is accepted) it falls back to a judge. Voice-native efforts each cover part of the space: live adaptive callers, native S2S evaluation, stateful tools, repeated-run reliability. To our knowledge none of them treat multi-agent topology, handoff verification, and conjunctive verification as one job.

BenchmarkMulti-agent topologyHandoff verificationConjunctive verificationLive adaptive voiceNative S2SMulti-industryStateful toolsFinal-state verifierTool-adherence verificationRepeated-run reliability
MIVAS BenchYesYesYesYesYesYesYesYesYesYes
EVANoNoNoYesYesYesYesYesPartialYes
τ-VoiceNoNoNoYesYesYesYesYesNoPartial
VAmoS BenchNoNoNoYesYesNoYesNoPartialYes
Full-Duplex-Bench v3NoNoNoNoYesPartialPartialNoYesNo
VoiceAgentBenchNoNoNoNoNoPartialNoNoPartialNo

Yes · Partial · No. Source: mivas-bench README benchmark comparison.

Table 1 Coverage of voice-agent benchmarks across the ten properties MIVAS was designed around. MIVAS is the only row that stays yes across topology, handoff, and conjunctive checks.

Our reading of that landscape is that the missing pieces are not independent. Multi-agent topology is what makes handoff verification meaningful, and handoff verification is what makes a deterministic reward defensible on read-based tasks. MIVAS was designed around that dependency.

03 / Environment design

Harnesses as multi-agent voice systems

Our first contribution is representing each MIVAS environment as a multi-agent voice system. Each subagent represents a different department or qualification stage. Responsibilities are mutually exclusive between subagents, with a small set of global session actions (escalate to a human, end the call) available from every node.

Patterns in multi-agent architectures: handoffs, supervisor, parallel, sequential, router, loop, aggregator, network, hierarchical
Figure 1 Common multi-agent patterns. MIVAS adopts the handoff pattern, which is what production voice deployments actually look like: one specialist transfers the caller to the next, and the transfer is invisible to the caller.

This is also the pattern that ElevenLabs and Bland promote through workflow-based agent development, and in our experience it is what survives contact with production. A single prompt carrying thirty tools degrades quickly. A graph of narrow specialists, each with a handful of tools, holds up.

Harness and pack. We split the repository so that the model runtime and the industry world can be paired without rewriting either side. The harness is the provider adapter: it carries bidirectional audio, builds the agents the blueprint declares, performs native handoffs, and POSTs industry tools. It owns no policy, prices, or schema. The pack owns the multi-agent blueprint, production-length prompts, tool definitions, database schema and seed, and the task suite. agent_blueprint.json is the interface. The same OpenAI Realtime harness runs Straus, Halverson & Reed, or Kestrel; the same Kestrel pack runs against Gemini Live or Grok. A scored environment is always one harness plus one pack.

Three industries. We wrote three complete companies from production voice lines we already knew, then replaced every name, chart, and order so the world is fictional, versioned, and shared across harnesses. Each pack exposes a tool server with fully implemented functions: lookups read real rows, writes mutate real rows, and every tool has a locked argument schema.

IndustryReplica companyInbound lineGraphDistinctive constraint
HealthcareStraus DermatologyRobin7-node DAGIdentity is the PHI gate before any protected desk
LegalHalverson & ReedHal5-node chainConflict screening before any facts of the matter
Customer supportKestrel ElectronicsKestrel7 instruction setsOrder-bound identity gate; fraud desk intentionally ungated

Table 2 The three launch packs. Full fixtures, prompts, and tools are on the industries page.

Straus is the representative case. Reception greets, answers public office facts through list_locations, and routes. Identity is the PHI gate: name and date of birth, then verify_identity before any protected desk. Scheduling, coverage, cosmetic, billing, and clinical each own one desk. Scheduling and cosmetic are sinks. Billing and clinical may hop forward to scheduling and never back to identity. Escalation is one global tool, transfer_to_human.

Straus Dermatology handoff graph: inbound call to reception, then identity, coverage, scheduling, cosmetic, clinical, billing, human, and call ends
Figure 2 The Straus Dermatology handoff graph. Solid edges are specialist transfers; dotted edges escalate to a human. The scheduling node cannot see billing tools, and billing cannot book an appointment.

Isolated state. Every conversation is keyed to its Bluejay simulation result id. The first tool call copies schema.sql and seed.sql into a SQLite file for that call alone, and later tools reopen it. Harnesses are dumb pipes: they POST to the tool server and return the envelope. This is the property that lets us diff a database at hangup and attribute every write to exactly one conversation.

The harness is part of the system under test. Providers do not expose the same multi-agent primitive, so the same DAG is implemented three ways. OpenAI Realtime and our cascaded baseline use a native agent graph. Grok and Qwen keep one WebSocket and swap instructions with session.update. Nova Sonic and Gemini Live open a new stream per specialist and seed it so the next voice continues mid-call. We treat those adapters as part of what is being measured, because a production voice agent is the model plus the runtime that can actually host the graph. Those details are also why two S2S models with similar published capability can diverge on the same locked case. See Model Harnesses.

04 / Task construction

Locked callers, locked tasks

A world without a specified caller is a demo. Bluejay Digital Humans are locked users: an identity, one objective, a trait profile, and scripted replies that keep the measurement on the task. They are live, adaptive voice callers rather than IVR script-readers. They ask again when an office-level answer is missing, and they decline a waitlist, a person, or a confirmation text when the case forbids it. Creativity is pinned so avoidable variance does not wash out the score.

We locked 72 cases per industry: 60 base tasks and 12 audio variants that replay selected cases under cafe or office noise, or a degraded signal, with intent, tools, and expected state held constant. Difficulty is balanced at 24 easy, 24 medium, and 24 hard across five workflow categories plus a sixth category for regulation, refusals, escalation, and impersonation. Every task.json states the caller, the objective, the expected handoff path, the required tools with constrained arguments, and the exact final database. The tasks page is the public lockfile, and every case was replayed against fresh seed state before release.

05 / Verification

A deterministic conjunctive verifier

Our second contribution is the verifier. MIVAS scores every conversation with three independent deterministic checks and defines combined pass as their conjunction.

  1. 01Database-state adherence. The hangup snapshot of write-bearing tables must match the expected state produced from the same seed. Catalog rows, timestamps, and unconstrained prose are ignored.
  2. 02Tool-call adherence. Required industry tools must fire with the constrained arguments the task locks and return the expected output. Extra industry tools are permitted.
  3. 03Handoff adherence. The expected specialist transfers must appear as an in-order subsequence of the transfer tools the call made. Extra hops are permitted. An empty expected path always passes, which is how some refusals stay reception-only.
Combined pass: tools AND handoff AND final DB equals pass
Figure 3 Combined pass is the logical AND of the three gates on the same conversation, not a weighted score.

The read-based case is the reason tool adherence cannot be folded into final state. Healthcare C1-E1 is our teaching example. Dana Whitfield is a new patient who wants a yes or no on whether Aetna is accepted at the Park Avenue office, not at the practice in general. If the agent says it takes Aetna without naming the office, she asks which office was checked. The required tool is check_plan_accepted with carrier=aetna and location_id=loc_park_ave. The expected hop is coverage. The expected database is the unchanged seed. A state-only verifier passes an agent that answered from memory. Ours does not.

Handoff adherence catches a third failure class that neither of the other gates can see: tools firing from the wrong node. If reception books the visit, or coverage writes a field it should never touch, the arguments may match and the final tables may match, but the specialist graph has leaked. We classify that as a scoping failure: the runtime allowed an agent to act outside its role.

We continue to record transcript quality, latency, cost, and a small set of 1 to 5 judge scores on every export, because judges catch spoken-policy failures the gates cannot see, such as a correct check_plan_accepted whose result is never said aloud. Those are diagnostics. None of them enter combined pass. The verifier contract is specified in full in the methodology.

06 / Reward shaping

Optimal paths and partial credit

Each MIVAS task has an optimal solution: database state matches, required tools fire, and the handoff path is the shortest route through the graph that satisfies the task. Handoff adherence is the component that makes partial reward well-defined. Any path that produced the correct final state with correct tool calls is rewarded. Any path longer than optimal receives a partial reward. Full reward requires the optimal path, the expected state, and the expected tool calls with parameters and output.

Figure 4 shows a real failure of this kind. The task expected identity, then coverage, then scheduling. The model verified identity and transferred directly to scheduling, skipping the coverage check that decides whether the visit is covered at all. The booking may look correct in isolation. The route was not.

Expected path transfer_to_identity, transfer_to_coverage, transfer_to_scheduling versus actual path transfer_to_identity, transfer_to_scheduling
Figure 4 Expected versus actual handoff path on the Straus graph. The skipped coverage hop fails handoff adherence even when the final database matches.

We think this gradient matters more than any single pass rate. A binary verifier collapses “reached the right state eventually” and “reached it cleanly” into the same label, and an RL loop trained on that label learns nothing about routing. Handoff adherence separates them deterministically, without a judge, which is what makes MIVAS tasks usable as RLVR environments rather than only as an evaluation.

Every scored conversation retains its evidence: the recording, the trace with every tool span and handoff on a timeline, and the hangup snapshot. A failed gate can be traced to the moment in the audio where the route diverged.

Recording and trace timeline for a MIVAS conversation showing audio, agent turns, tool spans, and cost
Figure 5 A scored conversation in Bluejay: audio track, agent turns, tool spans, and handoff markers on one timeline, with per-conversation cost.

07 / Reliability

Pass@1 versus pass^5

A production voice line is judged on whether it can do the same request again, so every locked case runs five times, independently, each with its own result id and its own SQLite file. We report two statistics on the leaderboard. Pass@1 is the single-attempt rate: scored conversations that cleared all three gates, among pass and fail only. Pass5 estimates the probability that the same agent succeeds on all five attempts of a task, using the unbiased estimator from tau-bench (Yao et al., 2024): C(c, 5) / C(n, 5) per task, then the unweighted mean across tasks. With five trials that is 100% only when all five pass. Tasks with fewer than five scored trials are left out of the mean.

How pass@1 and pass5 diverge

Each row is one locked task with five scored trials. pass@1 is c / 5. pass5 is C(c, 5) / C(5, 5): 100% only when all five trials pass. The published pack score is the unweighted mean of those per-task pass5 values.

5/5 passed

pass@1100.0%
pass^5100.0%

4/5 passed

pass@180.0%
pass^50.0%

3/5 passed

pass@160.0%
pass^50.0%

2/5 passed

pass@140.0%
pass^50.0%

1/5 passed

pass@120.0%
pass^50.0%

0/5 passed

pass@10.0%
pass^50.0%

The gap between the two is the reliability claim. Three passes in five is 60% pass@1 and 7.8% pass5. HumanEval-style pass@k asks whether at least one of k samples succeeds. We report passk because a front desk that books the visit once in five calls is not a front desk.

Scored runtimes

12

Industry packs

3

Locked tasks

216

Scored conversations

12,960

Exports keep one row per conversation with component passes, state differences, transcript, trace, latency, and estimated cost, so every leaderboard figure can be recomputed from evidence.

08 / Limitations

Where MIVAS stops today

MIVAS is built to become a single source of truth for voice AI performance across the economy. The launch release is not that yet, and the gaps are specific.

  • Three industries are not the economy. Healthcare, legal, and customer support are where voice AI is being adopted fastest, but a cross-industry score over three packs does not generalize to sectors we have not modeled. We are adding industries every week to widen that coverage and move MIVAS closer to a true cross-sector indicator.
  • One replica per industry. Each pack is a single fictional company with its own policies, fee math, and gates. A score on Straus Dermatology is a score on that practice’s rules, not on every dermatology group. Additional replicas within an industry are the natural next step once the industry list is broader.
  • Architectures vary within an industry. Not every production voice agent is a handoff graph. Some teams ship a single large prompt, a router, or a supervisor pattern. We centered MIVAS on multi-agent handoffs because that is the architecture we have seen succeed most often and adopt most widely, and because it is the one that makes handoff adherence measurable. Systems built another way are still measured on tools and final state, but the handoff gate reflects a design choice they may not have made.
  • US businesses, English callers. All three launch packs are US companies serving English-speaking callers. Other markets, languages, and accents are out of scope for this release.
  • Eight runtimes at launch. Providers ship new speech-to-speech models and handoff primitives frequently. The leaderboard is a snapshot of the runtimes and revisions we have evaluated, and it will move as harnesses are added and re-run.

09 / Key takeaways

What we learned building it

01

Production voice agents are specialist graphs. A benchmark that hosts the model in one prompt is measuring a system nobody deploys.

02

No single artifact verifies a call. Database state misses reads, tool traces miss wrong specialists, and judges are not deterministic. Conjunction of three deterministic gates covers what each misses.

03

Handoff adherence is the key component. It makes read-based tasks verifiable without a judge and makes partial reward well-defined. We consider it the most important source of training signal in the bench.

04

Reliability is a separate axis from capability. Pass5 exposes agents that succeed often but not repeatably.

05

The harness is part of the system under test. Provider-native handoff primitives differ enough to move scores on identical locked cases.

10 / Availability

Origins, code, and citation

MIVAS grew out of our time evaluating voice AI agents at Bluejay, from Fortune 10 deployments to solo developers. We kept seeing the same failure shapes: the right answer on the wrong specialist, the right route with a missing write, the correct outcome reached in twice the hops it should have taken. The bench is designed to measure exactly those. New harnesses start on a control industry that is never scored: receive a call, build the declared agents, complete a native handoff, invoke a tool, persist an appointment. Only then are they eligible for healthcare, legal, or customer support.

MIVAS Bench README on GitHub with links to the leaderboard, technical blog, Hugging Face dataset, and industries
Figure 6 The mivas-bench repository. Leaderboard, methodology, dataset, and industry packs are linked from the README.

The code, packs, task lockfiles, and exports are at github.com/bluejay-ai-dev/mivas-bench. Contributions of new harnesses and new industries are welcome as pull requests. Bluejay Labs exists to accelerate trustworthy voice AI adoption, and MIVAS is our first step. Results and future work are published here on Labs.

Bluejay Labs: Accelerating Trustworthy Voice AI Adoption
Figure 7 Bluejay Labs. Our first benchmark, MIVAS, indexes speech-to-speech model performance across economic sectors.

How to cite

@misc{siddiqi2026mivasbench,
      title={MIVAS Bench: Multi-Industry Voice Agent Simulation Bench},
      author={Faraz Siddiqi and Yash Savalia},
      year={2026},
      url={https://github.com/bluejay-ai-dev/mivas-bench},
}