Bluejay / Labs
Open navigation

Working paper / 2026

MIVAS

Multi-Industry Voice Agent Simulation Bench

Abstract

Multi-Industry Voice Agent Simulation Bench (MIVAS) is an indicator of voice AI performance across economic sectors. MIVAS is comprised of tasks and verifiers with granular rewards executed in production-grade multi-agent RL environments across industries where voice AI is being adopted the fastest. Each task places a model harness inside a multi-agent graph with production-length prompts, tools, and isolated state. Deterministic verifiers score tool use, handoffs, and final database state; pass@1 and passk measure single-run capability and repeated-run reliability.

GitHub repository ↗

Clone the repository, pair a model harness with an industry task suite, and reproduce the evaluation environment.

Back to Model Harnesses

Gemini 3.8 Live Extended Thinking

Speech-to-speech · gemini-3.8-live-extended-thinking

What the model is

Google gemini-3.8-live-extended-thinking, the 3.8 Live model with a reasoning phase before each spoken turn (thinking_level HIGH). Native audio in and out on the Gemini Live socket. The harness is the same LiveKit Agents worker as the other Gemini runtimes, selected by GEMINI_LIVE_MODEL, so the two 3.8 rows share one image and differ only in that variable and the thinking level.

Provider model card · Harness in mivas-bench

How the harness was built

The adapter lives in voice-agent-harnesses/gemini/3.8-live/. Extended thinking rejects FunctionResponse.scheduling outright (the socket closes with 1007 on the first tool response), so the harness clears scheduling on every function response and lets the model answer tool results on its own. It refuses to connect without a thinking_level.

One caller turn becomes several generations (filler, tool call, spoken answer), each closed with turn_complete and an interaction status. livekit-plugins-google 1.8 raises a barge-in on every server-initiated generation and cuts the previous playout; the harness suppresses that for extended so the model's own multi-generation turns play out in full. Real barge-ins still arrive as server_content.interrupted.

If the model goes idle after a tool result without speaking, the harness sends one empty completed turn, at most once per caller turn, so the result is spoken instead of waiting for the caller to prompt it.

Multi-agent assumptions

Gemini Live cannot swap instructions or tools on an open socket. Each blueprint agent gets its own Live session with that stage's prompt and tools. A handoff copies the prior chat context onto the new stage and appends the tool outputs of the previous stage as a note, so the specialist sees the ids the receptionist looked up.

Every stage prompt ends with the runtime's pack clock: the one-line date plus a calendar block naming that date as the only current date. The block was added on 2026-09-28 after plain 3.8 Live was observed substituting the real-world date for the pack date once a session carried handoff history; extended searched the pack calendar on its own on most calls.

Runtime

The September 2026 evaluation is three five-run suites, 1,080 scored conversations, on healthcare (run 354876), legal (354815), and customer-support (354877) with a twelve minute cap and 24 concurrent calls per pack. Calls with no scorable conversation were rerun on the same simulation and filled in; every task has five scored trials.

Cost/hr is Gemini list price on traced per-turn token usage (ai.google.dev/gemini-api/docs/pricing, checked 2026-09-29): $0.75 text and $3.00 audio per 1M input tokens, $4.50 text and $12.00 audio per 1M output tokens, output including thinking tokens. Usage is read from the realtime_inference span the LiveKit plugin writes for each model turn, or its child realtime_metrics span when the packet landed late. 8,223 of 26,026 turns, almost all tool-call turns, never received a packet; each is billed from what the trace measured, its reply audio at Google's $0.018 per minute, new caller audio at $0.005 per minute, tool-call text at four characters per token, and the context the neighbouring packet was billed for, which adds 22.6% to the row. Thinking tokens emitted on those packet-less turns are not recoverable and are the remaining understatement.

Deployment

Each industry runs its own Kubernetes Deployment (mivas-gemini-3-8-live-extended-<industry>) as a LiveKit worker registered with LiveKit Cloud; the simulation engine dials the SIP host and a dispatch rule maps the SIP user to the agent name. No CHIRP ingress.

Local: export GOOGLE_API_KEY and the LiveKit trio, set GEMINI_LIVE_MODEL=gemini-3.8-live-extended-thinking and GEMINI_THINKING_LEVEL, then .venv/bin/python voice-agent-harnesses/gemini/3.8-live/agent.py start.

Source: mivas-bench, run.py, k8s/deployment.yaml.

Scores

pass@1 among pass/fail only. pass5 is reliability from the five-run healthcare, legal, and customer-support suites.

Industrypass@1pass5ToolsHandoffFinal DBLatencyCost/hr
Healthcarescored82.5%63.9%82.8%96.7%91.7%2.61s$6.26/hr
Legalscored80.0%45.8%83.1%93.9%81.7%2.29s$6.28/hr
Customer supportscored83.3%56.9%86.7%93.3%91.7%2.56s$6.12/hr
Cross-industry81.9%55.5%84.2%94.6%88.4%2.49s$6.22/hr