Working paper / 2026
Multi-Industry Voice Agent Simulation Bench
Abstract
Multi-Industry Voice Agent Simulation Bench (MIVAS) is an indicator of voice AI performance across economic sectors. MIVAS is comprised of tasks and verifiers with granular rewards executed in production-grade multi-agent RL environments across industries where voice AI is being adopted the fastest. Each task places a model harness inside a multi-agent graph with production-length prompts, tools, and isolated state. Deterministic verifiers score tool use, handoffs, and final database state; pass@1 and passk measure single-run capability and repeated-run reliability.
Clone the repository, pair a model harness with an industry task suite, and reproduce the evaluation environment.
MIVAS groups its evaluation tasks into healthcare, legal, and customer-support suites. Each suite is a fully specified production environment: a multi-agent graph, node-level prompts, domain tools, seeded database state, and locked simulated-user tasks.
The environment and verifier contract stay fixed across model harnesses. A trajectory is evaluated for required tool use, handoff routing, and final state, enabling controlled comparisons across domains without collapsing failures into a single undifferentiated score.
Dermatology front desk across seven specialists: access, identity, scheduling, coverage, cosmetic quotes, billing, and clinical triage.
Straus Dermatology
Plaintiff-firm intake under conflict-before-facts, no legal advice, and attorney-only declines.
Halverson & Reed
Retail support behind an order-bound identity gate: delivery, returns, TechCrew, membership, and an ungated fraud desk.
Kestrel Electronics