Bluejay / Labs
Open navigation

Working paper / 2026

MIVAS

Multi-Industry Voice Agent Simulation Bench

Abstract

Multi-Industry Voice Agent Simulation Bench (MIVAS) is an indicator of voice AI performance across economic sectors. MIVAS is comprised of tasks and verifiers with granular rewards executed in production-grade multi-agent RL environments across industries where voice AI is being adopted the fastest. Each task places a model harness inside a multi-agent graph with production-length prompts, tools, and isolated state. Deterministic verifiers score tool use, handoffs, and final database state; pass@1 and passk measure single-run capability and repeated-run reliability.

GitHub repository ↗

Clone the repository, pair a model harness with an industry task suite, and reproduce the evaluation environment.

Industries

MIVAS groups its evaluation tasks into healthcare, legal, and customer-support suites. Each suite is a fully specified production environment: a multi-agent graph, node-level prompts, domain tools, seeded database state, and locked simulated-user tasks.

The environment and verifier contract stay fixed across model harnesses. A trajectory is evaluated for required tool use, handoff routing, and final state, enabling controlled comparisons across domains without collapsing failures into a single undifferentiated score.