Bluejay / Labs
Open navigation

Working paper / 2026

MIVAS

Multi-Industry Voice Agent Simulation Bench

Abstract

Multi-Industry Voice Agent Simulation Bench (MIVAS) is an indicator of voice AI performance across economic sectors. MIVAS is comprised of tasks and verifiers with granular rewards executed in production-grade multi-agent RL environments across industries where voice AI is being adopted the fastest. Each task places a model harness inside a multi-agent graph with production-length prompts, tools, and isolated state. Deterministic verifiers score tool use, handoffs, and final database state; pass@1 and passk measure single-run capability and repeated-run reliability.

GitHub repository ↗

Clone the repository, pair a model harness with an industry task suite, and reproduce the evaluation environment.

Back to Model Harnesses

GPT Live 1 (sol-low)

Speech-to-speech · gpt-live-1 + gpt-5.6-sol

What the model is

OpenAI GPT Live 1 is a full-duplex voice model that listens and speaks at the same time while delegating reasoning, tool use, and actions to a backend Responses model. This variant (sol-low) pins the backend to gpt-5.6-sol with reasoning effort low. The voice layer and the image are the same for every GPT Live 1 variant; only the backend model and effort change.

Provider model card · Harness in mivas-bench

How the harness was built

The adapter lives in voice-agent-harnesses/openai/gpt-live-1/. live.py opens one wss://api.openai.com/v1/live/sessions session per call, starts it with session.start, and answers every delegated function call with response.item.create followed by one response.create. variants.json sets GPT_LIVE_BACKEND_MODEL=gpt-5.6-sol and GPT_LIVE_REASONING_EFFORT=low.

The simulation engine reaches the pod over CHIRP, a WebSocket transport carrying 16 kHz pcm_s16le. The session is opened at audio/pcm 16000 so audio passes through byte for byte with no resampling, gating, or mute; the model is full duplex and owns barge-in. Industry tools POST to {TOOL_SERVER_URL}/tools/{name} with X-Mivas-Call-Id. Backend tools are registered with strict set to false so optional arguments are never invented.

Until the caller's first audio frame arrives the harness feeds 100 ms silence frames, which starts the session clock so the pack's fixed greeting is spoken without waiting for the caller.

Multi-agent assumptions

The industry DAG maps onto soft handoffs on the one live session. A transfer tool triggers session.update, which swaps the backend's instructions and tools to the target stage, then a short session.instructions.append tells the voice layer its role changed. Only the target stage's tools are ever registered, and history stays on the service; nothing is replayed.

The pack owns greeting text and tool policy. end_call is answered locally; the backend finishes its closing turn, the voice layer speaks it, and the harness closes the session once that audio goes quiet.

Runtime

The September 2026 evaluation includes 1,072 scored conversations across three five-run industry suites (runs 354856, 354865, 354880). Eight void attempts, calls that never connected or were never evaluated by the platform, are excluded from pass rates.

Cost/hr is list price for the whole tested system: $0.05 per minute for the gpt-live-1 voice session on the API-reported session seconds, plus the traced gpt-5.6-sol token usage from every backend call (short-context standard rates, $4 input, $0.40 cached, $20 output per 1M, from the OpenAI pricing page, checked 2026-09-29). Calls the voice layer answered without delegating carry no backend tokens.

Deployment

Each industry runs its own Kubernetes Deployment, Service, Ingress, and Bluejay agent (openai-gpt-live-1-sol-low-<industry>). Bluejay digital humans called the deployed healthcare, legal, and customer-support environments as external users, 20 concurrent calls per industry. Every scored conversation started from isolated seeded state and was checked against the same deterministic gates as the other MIVAS harnesses.

Source: mivas-bench, run.py, k8s/deployment.yaml.

Scores

pass@1 among pass/fail only. pass5 is reliability from the five-run healthcare, legal, and customer-support suites.

Industrypass@1pass5ToolsHandoffFinal DBLatencyCost/hr
Healthcarescored90.5%78.3%90.5%97.5%96.6%1.74s$4.78/hr
Legalscored80.2%42.3%80.5%92.5%82.5%1.88s$4.60/hr
Customer supportscored85.7%70.6%86.5%93.8%92.1%1.72s$4.78/hr
Cross-industry85.5%63.7%85.8%94.6%90.4%1.78s$4.72/hr