Bluejay / Labs
Open navigation

Working paper / 2026

MIVAS

Multi-Industry Voice Agent Simulation Bench

Abstract

Multi-Industry Voice Agent Simulation Bench (MIVAS) is an indicator of voice AI performance across economic sectors. MIVAS is comprised of tasks and verifiers with granular rewards executed in production-grade multi-agent RL environments across industries where voice AI is being adopted the fastest. Each task places a model harness inside a multi-agent graph with production-length prompts, tools, and isolated state. Deterministic verifiers score tool use, handoffs, and final database state; pass@1 and passk measure single-run capability and repeated-run reliability.

GitHub repository ↗

Clone the repository, pair a model harness with an industry task suite, and reproduce the evaluation environment.

Technical note / 2026

The conjunctive verifier

00

Overview

We score every conversation with three independent deterministic checks, and take their conjunction as the result of the call. No judge model sits in that path.

Our second contribution with MIVAS is the verifier. The argument for why we built it this way is in the whitepaper and the methodology; this page is the contract those pieces point at.

The verifier works in three parts. The first is database-state adherence: at hangup we compare the write-bearing tables against the expected state produced from the same seed, and we ignore catalog rows, timestamps, and unconstrained prose. The second is tool-call adherence: the required industry tools have to fire with the constrained arguments the task locks and return the expected output, while extra industry tools are permitted. The third is handoff adherence: the expected specialist transfers have to appear as an in-order subsequence of the transfer tools the call actually made. Extra hops are permitted. An empty expected path always passes, which is how some refusals stay reception-only.

result = T and H and D

If there is no hangup dump to compare, D is not evaluated and the result is T and H. Tool-call adherence and handoff adherence are always evaluated. An empty expected tool list, or an empty expected handoff path, is a defined pass of that half. The gates do not read the transcript. Spoken-policy misses stay on auxiliary judges and do not enter combined pass.

01

Rationale

The read-based case is the reason tool adherence cannot be folded into final state, and the scoping case is the reason handoff adherence cannot be folded into either of the other two.

Failure classToolsHandoffDBWhy conjunction
Answered a read from memory; seed unchangedfailmay passpassτ-bench-style state-only scoring accepts this. C1-E1 is the teaching case.
Required tool fired with the wrong constrained argfailmay passmay passcheck_plan_accepted(carrier=aetna, location_id=loc_windermere) is not Park Avenue.
Correct tools and writes, from the wrong specialistpassfailpassReception booked the visit. Arguments match; the graph leaked.
Required hop skipped; later specialist still writesmay passfailmay passidentity → scheduling, skipping coverage. The booking can look correct.
Required hop present, extra desks in betweenmay passpassmay passBinary gate passes as in_order_with_extras. The verdict is the RLVR hook.
Correct route and tools; write missing or extrapasspassfailA spoken confirmation that never called book_appointment, or a write the lockfile forbids.
Healthcare send_sms fired when the case forbade itfailmay passmay passThe only forbidden extra. Other extra industry tools are allowed.

A database-state diff, the verifier τ-bench established for text agents, only sees what was written. Many of the locked cases in MIVAS are satisfied by reads: plan acceptance at a named office, a filing deadline, order status, eligibility. If the agent answers from memory or from a practice-wide guess, the seed is unchanged and a state-only verifier would pass. Healthcare C1-E1 is our teaching example. Dana Whitfield is a new patient who wants a yes or no on whether Aetna is accepted at the Park Avenue office, not at the practice in general. If the agent says it takes Aetna without naming the office, she asks which office was checked. The required tool is check_plan_accepted with carrier=aetna and location_id=loc_park_ave. The expected hop is coverage. The expected database is the unchanged seed. A state-only verifier passes an agent that answered from memory. Ours does not.

Handoff adherence catches a third failure class that neither of the other gates can see: tools firing from the wrong node. If reception books the visit, or coverage writes a field it should never touch, the arguments may match and the final tables may match, but the specialist graph has leaked. We classify that as a scoping failure. The runtime allowed an agent to act outside its role. A handoff is a session-level operation rather than an industry POST, so it may leave no row at all, which is why the database check cannot stand in for it.

02

Tool-call adherence

Every expected call has to match some unused actual of the same name. Extra industry tools are allowed, with one healthcare exception. We do not compare raw JSON.

Actuals of the same name arrive grouped, not in wall-clock order. We flatten them and sort by start_offset_ms when it is present. Handoff scoring depends on this. A billing transfer that happened to be grouped first still gets scored after identity once the timestamps are put back.

Matching is greedy and one-to-one. For each expected call, in lockfile order, we take the first unused actual of the same name that satisfies the matcher. A required name that never matches is missing.

T = (missing is empty) and (forbidden extra is empty)
score_T = |hit| / |expected|     # 1 if nothing was expected

The forbidden-extra set is empty except in healthcare: send_sms is required when the lockfile lists it and forbidden when it does not. Confirmation texts are scored, never optional. Everything else extra is fine, including end_call.

What we actually compare

Live arguments are projected onto the industry schema. Harness routing keys and undeclared keys are dropped. Empty values count as absent. If the expected call has no arguments, or the live call has none left after that projection, it is a name-only hit. We will not invent requirements the lockfile did not write. That is how a transfer recorded as just a name still matches a live transfer that only carried harness routing.

When both sides have industry arguments, each expected key is one of two things:

  • Fact, enum, id, number, boolean, date. Must be present and equal after aliases. A schema enum, pattern, date format, or a key ending in _id / _e164 is never treated as prose.
  • Prose. We ignore the wording. If a value is present, it just has to be non-empty. The agent can paraphrase notes. It cannot send a blank.

Fact keys, by industry

IndustryKeys
Healthcarefull_name, first_name, last_name, dob, zip, carrier, member_id, destination, queue, location_id, provider_id, slot_id, appointment_id, start, end, new_start, new_end, service_date, group_number, channel, template_id, medication_name, visit_class, topic, service, order_type, appointment_type_code, cancellation_reason_code, fee_line_item_id, line_item_id, next_intent, subscriber_relationship
Legalopposing_party, practice_area, state, incident_date, reason_code, slot_id, confirmation_token, matter_id, channel, for_whom, provider, evaluation_id, attorney_id, phone, earliest_date, reason
Customer supportorder_number, postal_code, card_last4, reason_code, rma_number, sku, competitor, competitor_price, in_stock, opened, service_type, fee_disclosed_acknowledged, proration_acknowledged, confirmation_token

Prose keys, by industry

IndustryKeys
Healthcarenotes, patient_safe_message; reason when it is not an enum
Legalsummary, note, message, handoff_summary, full_name (presence-only). reason is unconstrained even when listed as a fact key.
Customer supporttopic, fee, query, issue, full_name, phone, reason

Aliases, so the environment and the grader agree

Equality is not raw JSON. We apply the same aliases the tool server already accepts, so a call the environment honored cannot be scored as missing. Park Avenue is loc_park_ave; "Park Avenue office in Manhattan" is not, and fails. Location sets are exact after the same alias, so adding Windermere to a Park Avenue-only search fails. Carriers collapse to slugs (Blue Cross Blue Shield is bcbs). Phones compare digits. An expected any_of is an or.

Legal opposing_party and for_whom use token-level Levenshtein with a small phonetic fold, so Daniel Okonkwo matches Daniel Oconquo, while New cases does not satisfy new cases intake. Legal full_name is presence-only. Healthcare full_name is exact after casefold.

Output comparison is narrow on purpose. If both sides recorded ok or error_code, those have to match. Nested data is not compared. A correct check_plan_accepted with ok: true still hits, even if the lockfile also wrote data.accepted.

03

Handoff adherence

The expected specialist transfers have to show up, in order, as a subsequence of the transfers the call actually made. Extra hops are allowed. An empty expected path always passes.

Actual hops are the flattened, wall-clock-sorted tool names, kept only if they sit in the industry transfer set. end_call is not a hop. transfer_to_human is a hop only in healthcare. Legal escalates with escalate_to_human, which is not in the legal set, so it does not move the pointer.

Transfer tools scored as hops

IndustryTools
Healthcaretransfer_to_identity, transfer_to_scheduling, transfer_to_coverage, transfer_to_cosmetic, transfer_to_billing, transfer_to_clinical, transfer_to_human
Legaltransfer_to_screening, transfer_to_intake, transfer_to_scheduling, transfer_to_client_services
Customer supporttransfer_to_verification, transfer_to_orders, transfer_to_returns, transfer_to_service, transfer_to_membership, transfer_to_fraud

The algorithm is a single pointer. No edit distance, no alignment:

i = 0
for hop in actual_handoffs:
    if i < len(expected) and hop == expected[i]:
        i += 1
score_H = i / len(expected)      # 1 if expected is empty
H = (expected is empty) or (i == len(expected))

In-order subsequence

Expected [identity, coverage]. Actual [identity, billing, coverage]. The pointer only advances on an exact name match. Verdict: in_order_with_extras, binary pass, score 1.

  1. 01transfer_to_identitytransfer_to_identityi = 0 → 1match; advance the pointer
  2. 02transfer_to_billingtransfer_to_coveragei = 1extra hop; pointer stays
  3. 03transfer_to_coveragetransfer_to_coveragei = 1 → 2match; expected exhausted

Out of order fails. Expected [identity, billing] against actual [billing, identity] consumes nothing on the first hop and only identity on the second. i = 1, verdict incomplete. Repeating a hop does not help unless the lockfile repeated it.

VerdictBinaryscoreWhen
exacttrue1actual == expected, including both empty
in_order_with_extrastrue1every expected hop appeared in order; extra hops remain
none_requiredtrue1expected is empty and at least one transfer fired
incompletefalsei / |E|the pointer never consumed the full expected path

An empty expected path is how a reception-only refusal stays reception-only without punishing a later transfer the caller did not need. The binary gate does not punish a longer-than-optimal route: score_H is 1 as soon as every expected hop appeared. The difference between optimal and merely successful is the verdict, exact versus in_order_with_extras, and that is the hook for partial reward.

04

Database-state adherence

The hangup snapshot is compared to exp_db_state on write-bearing tables only, after a per-industry cleanup. Catalog rows are not part of the agent contract.

We keep the industry write tables, drop ignored keys, normalize phones and datetimes, and sort rows by canonical JSON so insertion order cannot fail a match.

Write-bearing tables

IndustryWrite tables
Healthcarepatients, appointments, waitlistpatients and appointments exact after canonicalization. Non-empty waitlist is a subset match; empty waitlist requires empty actual. Autogenerated waitlist ids are stripped.
Legalintakes, intake_notes, documents, holds, evaluations, messages, escalationsConstrained-field cover. Empty expected intakes and escalations still require empty actual. Empty expected messages, intake_notes, documents, holds, and evaluations allow extra live rows.
Customer supportcustomers, orders, order_items, order_cancellations, delivery_changes, membership_changes, holds, refunds, rmas, return_labels, price_matches, protection_plans, service_appointments, escalations, scam_reportsEmpty expected table requires empty actual. Non-empty expected table requires exact canonical equality.

Ignored fields

IndustryKeys
All industriescreated_at, description; catalog tables (locations, providers, products); tool_events
Legal extramessage, summary, note, slot_id, starts_at, attorney_id, datetime, caller_id, reason
Customer support extrareason, including reason nested inside holds.payload JSON

Datetimes collapse to minute precision; seconds and timezone are dropped, so 15:30 and 15:30:59.5 are the same instant. Remaining strings are stripped and casefolded. Phones compare national digits. Healthcare waitlist ids are stripped, because an autogenerated id would otherwise steal a later matching row. Empty legal fields are unconstrained: a blank expected incident_date accepts any live date, while practice_area still has to match.

Healthcare is exact on patients and appointments. Waitlist is a subset when the expected list is non-empty, and must be empty when we locked empty. Extra waitlist rows are legal on a waitlist task and illegal on a task that forbade them. Legal matches on the fields we actually constrained, not the whole row. Messages, notes, documents, holds, and evaluations may have extra live rows, including when expected is empty; intakes and escalations do not. Customer support is exact on every write table. Empty expected means empty actual. A reason the caller never stated is ignored, so the agent can fill rmas.reason without being graded on prose.

05

Worked example

Dana Whitfield wants a yes or no for Aetna at Park Avenue. This call transfers to coverage and writes nothing, so handoff and database pass. The plan check used the Windermere office. Tools fail. Combined fail.

C1-E1 · wrong office

Expected check_plan_accepted with location_id=loc_park_ave. Live call used loc_windermere. Two gates pass. Combined fail.

T · Tools · fail

Same tool name, wrong constrained argument. Park Avenue was required.

H · Handoff · pass

Reception transferred to coverage. The expected hop is present.

D · Database · pass

No writes. Hangup matches the seed.

result · Conjunction · fail

T and H and D is false. One miss fails the task.

Locked fields, abbreviated:

exp_handoff_path: ["transfer_to_coverage"]
exp_tool_calls:
  - { name: transfer_to_coverage }
  - { name: check_plan_accepted,
      parameters: { carrier: aetna, location_id: loc_park_ave } }
exp_db_state: unchanged seed

Healthcare C1-E1 is the teaching case from the rationale. She is a new patient. If the agent says it takes Aetna without naming the office, she asks which office was checked. She is not booking. The seed must not change.

This call hops reception → coverage and leaves the database alone. Handoff is exact. Database matches the seed. The live plan check is check_plan_accepted(carrier=aetna, location_id=loc_windermere). That is the wrong constrained argument. Tools fail, so T and H and D is false. Two gates are not enough.

06

Partial rewards

Each MIVAS task has an optimal solution: the database matches, the required tools fire, and the handoff path is the shortest route through the graph that satisfies the task. Combined pass does not distinguish that route from a longer one that still hits every expected hop.

Full reward vs partial vs fail

Same locked case, C1-H1. Expected path coverage → scheduling. Tools and final state match in every row except where noted. Combined pass is still a yes/no. The reward is not.

What happenedActual pathVerdictPass?Reward

Optimal path

Shortest route that satisfies C1-H1. Full reward.

coverage → schedulingexactyes1

Correct, longer

Expected hops in order, extra desk first. Combined pass, discounted training reward.

identity → coverage → schedulingin_order_with_extrasyesdiscount

Skipped required hop

Coverage never ran. Pointer stays at 0. Combined fail.

schedulingincompleteno0

Reception booked it

Tools and DB can still look right. Handoff is the only gate that sees the leak.

(no transfer)incompleteno0

Combined pass is still a boolean. It answers the evaluation question: did this conversation clear the bar? A training run can ask a second question: did it clear the bar on the shortest route? If the path is exact, the training reward is 1. If every expected hop appeared, but extra desks sat in between, the training reward is a discount between 0 and 1, chosen by the training run. If combined pass is false, the training reward is 0. We do not publish a discount on the leaderboard. The public number is pass or fail.

r_eval  = 1 if T and H and D else 0

r_train = 1         if r_eval and verdict == exact
        = discount  if r_eval and verdict == in_order_with_extras
        = 0         if not r_eval

A binary verifier collapses “reached the right state eventually” and “reached it cleanly” into the same label, and an RL loop trained on that label learns nothing about routing. The verdict is what makes the tasks usable as RLVR environments, not only as an eval.