Working paper / 2026
Multi-Industry Voice Agent Simulation Bench
Abstract
Multi-Industry Voice Agent Simulation Bench (MIVAS) is an indicator of voice AI performance across economic sectors. MIVAS is comprised of tasks and verifiers with granular rewards executed in production-grade multi-agent RL environments across industries where voice AI is being adopted the fastest. Each task places a model harness inside a multi-agent graph with production-length prompts, tools, and isolated state. Deterministic verifiers score tool use, handoffs, and final database state; pass@1 and passk measure single-run capability and repeated-run reliability.
Clone the repository, pair a model harness with an industry task suite, and reproduce the evaluation environment.
A model harness is the policy adapter around a speech model: transport, session state, audio, turn-taking, and the tool and handoff bridge into a shared task environment. Every row runs the same locked domain suite and deterministic verifiers, isolating the behavior of the model and its runtime.
Every harness is a folder in voice-agent-harnesses/ in the mivas-bench repo. It loads the industry's agent_blueprint.json and turns that graph into a runnable multi-agent voice runtime for one provider. The work splits into four parts: a transport that carries the caller's audio, a session layer that speaks the provider's realtime protocol, an audio bridge that carries speech between the caller and the model, and a tool router that connects tool calls to the shared task environment.
The harness stays deliberately thin. Prompts, greeting text, tool schemas, and the handoff graph all come from the industry pack, and the same locked suites run on every row of the leaderboard. When two harnesses score differently on the same task, the difference comes from the model and its runtime.
One call through a harness
01 Dial. Bluejay places the call and the harness answers. The connection carries the metadata that ties this conversation to its own database, traces, and evals. The harness then opens a live session with the speech model, loaded with the initial agent's prompt and tools, and the agent greets the caller.
The blueprint marks every tool as function, handoff, or session, and each kind takes a different path through the harness. The split keeps the parts of the benchmark that must be identical fully identical, while the parts that genuinely differ between runtimes stay native.
Function tools do the real work: booking an appointment, looking up an order, updating a record. Every harness sends these through the same generic pipe to the task environment and reads the result back to the caller. No harness carries special-case code for any tool, so the environment behaves the same no matter which model is on the call.
Handoff tools move the conversation between agents, and every runtime has its own native way of making that switch. The harness uses whichever mechanism the platform provides and carries the conversation history along. How gracefully a runtime handles those switches is part of what the benchmark measures.
Session tools end the call. The harness lets the goodbye finish, then closes the connection.
Each industry ships a seeded database behind a small state API, and both live inside the same pod as the harness. When a new conversation starts, the environment creates a fresh copy of the seeded database just for that call. Parallel conversations never share state, and nothing from one call can leak into another.
That choice is what makes deterministic verification possible. Every call starts from the same known state, so the verifiers can diff exactly the rows one conversation touched and compare them, along with the tools used and the handoff path, against what the task required. At hangup the final state is frozen and stored, and the evals read that frozen copy.
Each harness and industry pair is packaged as one container: the voice runtime, the industry's tool server, and its seeded database travel together. Scored runs deploy those containers to Kubernetes, one deployment per pair, and Bluejay's digital humans dial each one from the outside like real callers.
Packing everything into one unit keeps the benchmark reproducible. A pair can be rebuilt, redeployed, or scaled without touching any other pair, and a scored run depends on nothing outside its own container except the model provider's API. Build scripts, cluster configuration, and the rest of the operational detail live in mivas-bench.
Every scored harness, with the model it runs and its hourly cost. Open a row for the full writeup: the model, the harness around it, the multi-agent wiring, and how it is deployed.
GPT Live 1 (astra-medium)
$6.72/hr
OpenAI · gpt-live-1 + gpt-6-astra · openai/gpt-live-1@astra-medium
OpenAI's GPT Live 1 full-duplex voice layer delegating to a gpt-6-astra backend at medium reasoning effort.
GPT Live 1 (sol-low)
$4.78/hr
OpenAI · gpt-live-1 + gpt-5.6-sol · openai/gpt-live-1@sol-low
OpenAI's GPT Live 1 full-duplex voice layer delegating to a gpt-5.6-sol backend at low reasoning effort.
Gemini 3.8 Live Extended Thinking
$6.26/hr
Google · gemini-3.8-live-extended-thinking · gemini/3.8-live@extended
Google's Gemini 3.8 Live speech-to-speech model with extended thinking, run through the Gemini Live API.
Gemini 3.8 Live
$3.00/hr
Google · gemini-3.8-live · gemini/3.8-live
Google's Gemini 3.8 Live speech-to-speech model, run through the Gemini Live API.
OpenAI Realtime 2.1
$5.27/hr
OpenAI · gpt-realtime-2.1 · openai/realtime-2.1
OpenAI's gpt-realtime-2.1 speech-to-speech model, run through the OpenAI Realtime Agents SDK.
Grok Voice 2
$5.77/hr
xAI · grok-voice · grok/voice
xAI's grok-voice speech-to-speech model, run through the xAI Realtime WebSocket API.
Qwen Audio Realtime
$1.82/hr
Alibaba · qwen-audio-3.0-realtime-plus · qwen/audio-realtime
Alibaba's Qwen Audio Realtime speech-to-speech model, run through the DashScope realtime API.
Amazon Nova Sonic 2
$0.820/hr
Amazon · amazon.nova-sonic-v2 · aws/nova-sonic-2
Amazon's Nova 2 Sonic speech-to-speech model, run through Bedrock's bidirectional streaming API.
Gemini Flash Live 3.1
$2.78/hr
Google · gemini-3.1-flash-live · gemini/flash-live-3.1
Google's Gemini Flash Live 3.1 speech-to-speech model, run through the Gemini Live API.
Cascaded Baseline
$2.60/hr
Cascaded · flux-general-en + gpt-4.1 + eleven_flash_v2_5 · cascaded
A cascaded pipeline of separate speech models: Deepgram Flux for transcription, GPT-4.1 for reasoning, ElevenLabs Flash 2.5 for speech.
OpenAI Realtime 2.1 Mini
$1.36/hr
OpenAI · gpt-realtime-2.1-mini · openai/realtime-2.1-mini
OpenAI's gpt-realtime-2.1-mini speech-to-speech model, run through the OpenAI Realtime Agents SDK.
Gemini 2.5 Flash Native Audio
$1.25/hr
Google · gemini-2.5-flash-native-audio · gemini/2.5-flash-native-audio
Google's Gemini 2.5 Flash Native Audio speech-to-speech model, run through the Gemini Live API.