Bluejay / Labs
Open navigation

Working paper / 2026

MIVAS

Multi-Industry Voice Agent Simulation Bench

Abstract

Multi-Industry Voice Agent Simulation Bench (MIVAS) is an indicator of voice AI performance across economic sectors. MIVAS is comprised of tasks and verifiers with granular rewards executed in production-grade multi-agent RL environments across industries where voice AI is being adopted the fastest. Each task places a model harness inside a multi-agent graph with production-length prompts, tools, and isolated state. Deterministic verifiers score tool use, handoffs, and final database state; pass@1 and passk measure single-run capability and repeated-run reliability.

GitHub repository ↗

Clone the repository, pair a model harness with an industry task suite, and reproduce the evaluation environment.

Back to Model Harnesses

Gemini 3.8 Live

Speech-to-speech · gemini-3.8-live

What the model is

Google gemini-3.8-live, Gemini 3.8 Live without the reasoning phase. Native audio on the Gemini Live socket, answered turn by turn. Same LiveKit Agents worker and image as the extended-thinking row, selected by GEMINI_LIVE_MODEL.

Provider model card · Harness in mivas-bench

How the harness was built

The adapter lives in voice-agent-harnesses/gemini/3.8-live/. Tool responses are sent with FunctionResponseScheduling.INTERRUPT, because the default WHEN_IDLE never fires on a continuous SIP stream. High start and end of speech sensitivity with 500 ms silence so quiet telephone-band confirmations open a turn.

The greeting is pinned into the connect-time instructions and started with a realtime text kick, the proven speak-first path for the Gemini family.

Multi-agent assumptions

One Live session per blueprint agent, handoff by copying the chat context onto the new stage plus a note carrying the previous stage's tool outputs. Tool JSON Schema is flattened to the keys Gemini Live accepts.

Every stage prompt ends with the runtime's pack clock, the one-line date plus a calendar block. Plain 3.8 Live honours the one-line date in a fresh session but reverted to its real-world date once a session carried handoff history, which cost it 50 of 59 failed legal booking calls on the first run; the legal row is the rerun with the calendar block in place, and healthcare and customer-support, where the effect touched at most three conversations, are the original runs.

Runtime

The September 2026 evaluation is three five-run suites, 1,080 scored conversations, on healthcare (run 354868), legal (356269, the calendar-block rerun), and customer-support (354866) with a twelve minute cap and 24 concurrent calls per pack. Calls with no scorable conversation were rerun on the same simulation and filled in; every task has five scored trials.

Cost/hr is Gemini list price on traced per-turn token usage (ai.google.dev/gemini-api/docs/pricing, checked 2026-09-29): $0.75 text and $3.00 audio per 1M input tokens, $4.50 text and $12.00 audio per 1M output tokens. Usage is read from the realtime_inference span the LiveKit plugin writes for each model turn, or its child realtime_metrics span when the packet landed late. 140 of 13,562 turns (1%), mostly the final farewell whose socket closes before the packet lands, are billed from their measured reply audio at Google's $0.018 per minute plus the re-sent context of the neighbouring packet.

Deployment

Each industry runs its own Kubernetes Deployment (mivas-gemini-3-8-live-<industry>) as a LiveKit worker registered with LiveKit Cloud; the simulation engine dials the SIP host and a dispatch rule maps the SIP user to the agent name. No CHIRP ingress.

Local: export GOOGLE_API_KEY and the LiveKit trio, then .venv/bin/python voice-agent-harnesses/gemini/3.8-live/agent.py start. --check loads the blueprint without LiveKit.

Source: mivas-bench, run.py, k8s/deployment.yaml.

Scores

pass@1 among pass/fail only. pass5 is reliability from the five-run healthcare, legal, and customer-support suites.

Industrypass@1pass5ToolsHandoffFinal DBLatencyCost/hr
Healthcarescored70.0%48.6%70.3%95.0%86.7%2.54s$3.00/hr
Legalscored68.9%25.0%73.9%93.9%75.0%2.44s$2.81/hr
Customer supportscored79.2%56.9%80.0%94.7%91.9%2.46s$2.81/hr
Cross-industry72.7%43.5%74.7%94.5%84.5%2.48s$2.87/hr