MIVAS Speech to Speech Benchmark
Cross-industry pass5 · mean of (c/n)5 over locked tasks · higher is better
Published Sep 26, 2026
Working paper / 2026
Multi-Industry Voice Agent Simulation Bench
Abstract
Multi-Industry Voice Agent Simulation Bench (MIVAS) is an indicator of voice AI performance across economic sectors. MIVAS is comprised of tasks and verifiers with granular rewards executed in production-grade multi-agent RL environments across industries where voice AI is being adopted the fastest. Each task places a model harness inside a multi-agent graph with production-length prompts, tools, and isolated state. Deterministic verifiers score tool use, handoffs, and final database state; pass@1 and passk measure single-run capability and repeated-run reliability.
Clone the repository, pair a model harness with an industry task suite, and reproduce the evaluation environment.
Each row is a model harness on the same locked task suite and deterministic verifiers. pass@1 is single-run success, pass5 is five-run reliability, and cross-industry values are an unweighted mean of the published suites. Pick a scope once and every chart and table follows it.
Cross-industry pass5 · mean of (c/n)5 over locked tasks · higher is better
Published Sep 26, 2026
Frontiers
The staircase joins the harnesses nothing else beats on both, at cross-industry scope.
Cost versus latency
4 of 12 harnesses sit on the frontier: Nova Sonic 2 ($0.750/hr, 2.54s), Realtime 2.1 Mini ($1.31/hr, 2.48s), Qwen Audio ($1.67/hr, 2.25s), and GPT Live 1 (sol-low) ($4.72/hr, 1.78s). Nova Sonic 2 (26.9%) and Realtime 2.1 Mini (16.6%) make it on price alone with pass5 under 30%.
Latency versus accuracy
Only GPT Live 1 (sol-low) (1.78s, 63.7%) and GPT Live 1 (astra-med) (1.81s, 66.3%) are on the frontier; everything else is slower and less reliable than at least one of them. Between them the trade is 37 ms of median latency for 2.6 points of pass5.
Cost versus accuracy
5 of 12 harnesses form the frontier: Nova Sonic 2 ($0.750/hr, 26.9%), Qwen Audio ($1.67/hr, 31.4%), Gemini 3.8 Live ($2.87/hr, 43.5%), GPT Live 1 (sol-low) ($4.72/hr, 63.7%), and GPT Live 1 (astra-med) ($6.46/hr, 66.3%). Best step: Gemini 3.8 Live → GPT Live 1 (sol-low), 20.2 points of pass5 for $1.85/hr more; priciest: GPT Live 1 (sol-low) → GPT Live 1 (astra-med), 2.6 points for $1.74/hr.
Rankings
Best first, at cross-industry scope.
Latency
GPT Live 1 (sol-low) is fastest at a 1.78s median; 2.5 Flash Native is slowest at 3.89s, 2.2× as long. Two harnesses answer under two seconds (GPT Live 1 (sol-low) and GPT Live 1 (astra-med)), 8 sit between two and three, and 2 take three or more.
Cost per hour
List price runs from $0.750/hr for Nova Sonic 2 to $6.46/hr for GPT Live 1 (astra-med), a 8.6× spread. The cheapest route to a pass5 above 60% is GPT Live 1 (sol-low) at $4.72/hr, above 50% is GPT Live 1 (sol-low) at $4.72/hr, and above 40% is Gemini 3.8 Live at $2.87/hr.
pass5 by pack
Averaged over the 12 harnesses, pass5 is 44.5% in healthcare, 42.4% in customer support, and 23.8% in legal. Pack leaders: GPT Live 1 (sol-low) in healthcare (78.3%), GPT Live 1 (astra-med) and Realtime 2.1 in legal (50.0%), and GPT Live 1 (astra-med) in customer support (71.4%).
Table
Cross-industry · 12 harnesses
| 1 | 66.3% | E 92.0% M 88.2% H 83.2% | 87.8% | 88.1% | 95.0% | 92.3% | 1.81s | $6.46/hr | |
| 2 | GPT Live 1 (sol-low) OpenAI | 63.7% | E 92.8% M 84.4% H 79.2% | 85.5% | 85.8% | 94.6% | 90.4% | 1.78s | $4.72/hr |
| 3 | 55.5% | E 93.9% M 78.3% H 73.6% | 81.9% | 84.2% | 94.6% | 88.4% | 2.49s | $6.22/hr | |
| 4 | Gemini 3.8 Live Google | 43.5% | E 89.2% M 74.2% H 54.7% | 72.7% | 74.7% | 94.5% | 84.5% | 2.48s | $2.87/hr |
| 5 | OpenAI Realtime 2.1 OpenAI | 42.1% | E 88.0% M 60.6% H 38.1% | 62.2% | 71.6% | 88.6% | 75.8% | 2.05s | $5.11/hr |
| 6 | Grok Voice 2 xAI | 34.7% | E 89.7% M 51.1% H 17.2% | 52.7% | 60.8% | 88.4% | 70.2% | 2.70s | $6.40/hr |
| 7 | Qwen Audio Realtime Alibaba | 31.4% | E 83.9% M 44.9% H 23.9% | 50.9% | 57.2% | 82.5% | 70.1% | 2.25s | $1.67/hr |
| 8 | Amazon Nova Sonic 2 Amazon | 26.9% | E 86.4% M 37.8% H 12.5% | 45.6% | 52.3% | 79.7% | 63.4% | 2.54s | $0.750/hr |
| 9 | Gemini Flash Live 3.1 Google | 25.5% | E 84.5% M 36.4% H 20.0% | 47.0% | 49.6% | 86.7% | 72.2% | 2.36s | $2.42/hr |
| 10 | Cascaded Baseline Cascaded | 24.5% | E 75.3% M 20.5% H 11.7% | 35.8% | 38.1% | 70.8% | 65.2% | 3.10s | $2.30/hr |
| 11 | OpenAI Realtime 2.1 Mini OpenAI | 16.6% | E 69.2% M 29.0% H 10.2% | 36.3% | 42.2% | 71.5% | 56.8% | 2.48s | $1.31/hr |
| 12 | 12.0% | E 59.7% M 20.3% H 5.6% | 28.5% | 30.3% | 61.2% | 55.0% | 3.89s | $1.06/hr |
† Cost/hr figures marked on the charts carry a pricing caveat: