MIVAS Bench is an indicator of voice AI performance across economic sectors. It is a set of tasks and deterministic verifiers, with granular rewards, executed inside production-grade multi-agent environments in the industries where voice AI is being adopted fastest. The initial release covers healthcare, legal, and customer support. For each, we built a voice agent for a fictitious business in extreme detail: a tool server with fully implemented functions, an attached database, and a multi-agent architecture whose tools and scenarios are drawn from real deployments in that industry.
MIVAS makes two contributions to voice benchmarks. The first is realism. Each environment is a graph of specialist subagents, the way enterprises actually build voice agents today, with production-length prompts, provider-native handoffs, and isolated per-call state. The second is verifiable reward. Every conversation is scored by three deterministic checks, database state, tool calls, and handoff path, combined with a logical AND, and the handoff check yields a partial reward that distinguishes an optimal route from a merely successful one. No judge model sits in the scoring path.
The figure at the top of this page is one task executing end to end:
- Simulated conversation. A Bluejay Digital Human, a locked caller with an identity, one objective, and scripted replies, talks to the agent over live audio.
- Multi-agent voice architecture. The agent is a directed acyclic graph of specialists. Each node exposes its own toolset; end call, transfer, and escalate are shared session tools available everywhere.
- Deterministic verification. At hangup, DB-state adherence, handoff adherence, and tool adherence are checked independently and combined with AND. One miss fails the task.
Each locked case runs five independent times. Pass@1 reports single-run capability and pass5 reports whether the same capability holds across all five trials. Native speech-to-speech models are the systems under test; a cascaded STT, LLM, and TTS stack is the baseline. The bench is open source and open data, and every task is packaged as an RLVR environment.
We believe this combination gives MIVAS the highest training signal of any voice benchmark available today. The sections that follow explain why: what existing benchmarks measure, how we designed the environments and tasks, how the verifier works, how the reward is shaped, and what the scores do and do not say.
Can a voice benchmark be production-realistic and fully deterministic at the same time, and can its reward be granular enough to train on?







