Bluejay / Labs
Open navigation

Working paper / 2026

MIVAS

Multi-Industry Voice Agent Simulation Bench

Abstract

Multi-Industry Voice Agent Simulation Bench (MIVAS) is an indicator of voice AI performance across economic sectors. MIVAS is comprised of tasks and verifiers with granular rewards executed in production-grade multi-agent RL environments across industries where voice AI is being adopted the fastest. Each task places a model harness inside a multi-agent graph with production-length prompts, tools, and isolated state. Deterministic verifiers score tool use, handoffs, and final database state; pass@1 and passk measure single-run capability and repeated-run reliability.

GitHub repository ↗

Clone the repository, pair a model harness with an industry task suite, and reproduce the evaluation environment.

Back to Model Harnesses

Gemini 2.5 Flash Native Audio

Speech-to-speech · gemini-2.5-flash-native-audio

What the model is

Google gemini-2.5-flash-native-audio, Gemini 2.5 Flash Native Audio. Same Gemini Live transport as Flash Live 3.1, different model id and a generate_reply path that 3.1 does not honor.

Harness in mivas-bench

How the harness was built

2.5-flash-native-audio/agent.py shares gemini/harness.py with Flash Live 3.1. LiveKit SIP inbound, Puck, en-US, FunctionResponseScheduling.INTERRUPT, schema flattening, one Live socket per blueprint agent, and the update_agent handoff are the shared class.

2.5 can generate_reply. The opening uses session.generate_reply(instructions=…) with the pack greeting. Handoff entry calls generate_reply() on the new Stage. Connect-time instructions are the pack prompt plus today's date. speak_first is not pinned into the system prompt.

Multi-agent assumptions

Tools remain fixed at connect, so each specialist is still a new Gemini Live socket with the prior chat_ctx copied over. The industry DAG is the same Robin / Hal / Kestrel graph. 2.5 is the unscripted path (scripted=False): handoff entry uses generate_reply instead of the realtime text kick 3.1 needs.

Schema flattening, additionalProperties stripping, and the barge-in-safe update_agent swap are inherited from the shared harness. max_tool_steps=16.

Runtime

Same process model as 3.1: one active Gemini Live call per worker process, job_count_load(1), THREAD executor, 15 s farewell wait, 30 s tool POST. Same Kubernetes box: 1000m / 1Gi / 3Gi. Same LiveKit Cloud SIP secret.

Deployment

Same LiveKit worker Deployment as Flash Live 3.1. run.py --harness gemini/2.5-flash-native-audio --apply. Agent name defaults to mivas-gemini-2-5-flash-native-audio. Local start is the 2.5 agent.py start entrypoint.

Source: mivas-bench, run.py, k8s/deployment.yaml.

Scores

pass@1 among pass/fail only. pass5 is reliability from the five-run healthcare, legal, and customer-support suites.

Industrypass@1pass5ToolsHandoffFinal DBLatencyCost/hr
Healthcarescored43.1%19.4%43.1%62.8%66.9%4.07s$1.25/hr
Legalscored16.7%0.0%17.8%59.2%27.8%3.95s$0.990/hr
Customer supportscored25.8%16.7%30.0%61.7%70.3%3.65s$0.930/hr
Cross-industry28.5%12.0%30.3%61.2%55.0%3.89s$1.06/hr