Bluejay / Labs
Open navigation

Working paper / 2026

MIVAS

Multi-Industry Voice Agent Simulation Bench

Abstract

Multi-Industry Voice Agent Simulation Bench (MIVAS) is an indicator of voice AI performance across economic sectors. MIVAS is comprised of tasks and verifiers with granular rewards executed in production-grade multi-agent RL environments across industries where voice AI is being adopted the fastest. Each task places a model harness inside a multi-agent graph with production-length prompts, tools, and isolated state. Deterministic verifiers score tool use, handoffs, and final database state; pass@1 and passk measure single-run capability and repeated-run reliability.

GitHub repository ↗

Clone the repository, pair a model harness with an industry task suite, and reproduce the evaluation environment.

Leaderboard

Each row is a model harness on the same locked task suite and deterministic verifiers. pass@1 is single-run success, pass5 is five-run reliability, and cross-industry values are an unweighted mean of the published suites. Pick a scope once and every chart and table follows it.

Scope

MIVAS Speech to Speech Benchmark

Cross-industry pass5 · mean of (c/n)5 over locked tasks · higher is better

Published Sep 26, 2026

706050403020100
bluejay-labs
66.3%
63.7%
55.5%
43.5%
42.1%
34.7%
31.4%
26.9%
25.5%
24.5%
16.6%
12.0%
GPT Live 1 Astra
GPT Live 1 Sol
Gemini 3.8 Extended
Gemini 3.8 Live
Realtime 2.1
Grok Voice 2
Qwen Audio Realtime
Nova Sonic 2
Flash Live 3.1
Cascaded Baseline
Realtime 2.1 Mini
2.5 Native Audio
MIVAS Speech to Speech Benchmarkbluejay-labs

Frontiers

Two metrics at a time

The staircase joins the harnesses nothing else beats on both, at cross-industry scope.

Cost versus latency

$/hr · s
1.5s2.0s2.5s3.0s3.5s4.0s$0$2$4$6$8GPT Live 1 (astra-med) · Cost per hour $6.46/hr · Median latency 1.81sGPT Live 1 (sol-low) · Cost per hour $4.72/hr · Median latency 1.78sGemini 3.8 Extended · Cost per hour $6.22/hr · Median latency 2.49s · † 32% of turns (tool calls) sent no usage packet; those are billed from measured audio seconds, tool text and re-sent context at Google's published rates (22.6% of the row). Thinking tokens on those turns are not recoverable.Gemini 3.8 Live · Cost per hour $2.87/hr · Median latency 2.48sRealtime 2.1 · Cost per hour $5.11/hr · Median latency 2.05sGrok Voice 2 · Cost per hour $6.40/hr · Median latency 2.70s · † $0.08 per audio minute plus xAI's published $0.004 per text input, one per tool result. Without the text-input line Grok is $4.80/hr.Qwen Audio · Cost per hour $1.67/hr · Median latency 2.25sNova Sonic 2 · Cost per hour $0.750/hr · Median latency 2.54sFlash Live 3.1 · Cost per hour $2.42/hr · Median latency 2.36s · † 22% of turns sent no usage packet; stage-final ones are billed from measured audio and re-sent context at Google's published rates (12% of the row), the rest are folded into the next packet by Gemini.Cascaded · Cost per hour $2.30/hr · Median latency 3.10sRealtime 2.1 Mini · Cost per hour $1.31/hr · Median latency 2.48s2.5 Flash Native · Cost per hour $1.06/hr · Median latency 3.89s · † 38% of turns sent no usage packet; those are billed from measured audio seconds (32 tokens/s), tool text and re-sent context at Google's published rates (33% of the row).Nova Sonic 2Realtime 2.1 MiniQwen AudioGPT Live 1 (sol-low)2.5 Flash Native †CascadedGrok Voice 2 †Gemini 3.8 Extended †Gemini 3.8 LiveRealtime 2.1COST PER HOUR · USD / HR · ← betterMEDIAN LATENCY · VOICE-TO-VOICE · ← better

Latency versus accuracy

s · pass5
0204060801.5s2.0s2.5s3.0s3.5s4.0sGPT Live 1 (astra-med) · Median latency 1.81s · pass^5 66.3%GPT Live 1 (sol-low) · Median latency 1.78s · pass^5 63.7%Gemini 3.8 Extended · Median latency 2.49s · pass^5 55.5%Gemini 3.8 Live · Median latency 2.48s · pass^5 43.5%Realtime 2.1 · Median latency 2.05s · pass^5 42.1%Grok Voice 2 · Median latency 2.70s · pass^5 34.7%Qwen Audio · Median latency 2.25s · pass^5 31.4%Nova Sonic 2 · Median latency 2.54s · pass^5 26.9%Flash Live 3.1 · Median latency 2.36s · pass^5 25.5%Cascaded · Median latency 3.10s · pass^5 24.5%Realtime 2.1 Mini · Median latency 2.48s · pass^5 16.6%2.5 Flash Native · Median latency 3.89s · pass^5 12.0%GPT Live 1 (astra-med)GPT Live 1 (sol-low)Gemini 3.8 ExtendedGemini 3.8 LiveRealtime 2.1Grok Voice 2Qwen AudioNova Sonic 2CascadedRealtime 2.1 Mini2.5 Flash NativeMEDIAN LATENCY · VOICE-TO-VOICE · ← betterPASS^5 · RELIABILITY, % · better →

Cost versus accuracy

$/hr · pass5
020406080$0$2$4$6$8GPT Live 1 (astra-med) · Cost per hour $6.46/hr · pass^5 66.3%GPT Live 1 (sol-low) · Cost per hour $4.72/hr · pass^5 63.7%Gemini 3.8 Extended · Cost per hour $6.22/hr · pass^5 55.5% · † 32% of turns (tool calls) sent no usage packet; those are billed from measured audio seconds, tool text and re-sent context at Google's published rates (22.6% of the row). Thinking tokens on those turns are not recoverable.Gemini 3.8 Live · Cost per hour $2.87/hr · pass^5 43.5%Realtime 2.1 · Cost per hour $5.11/hr · pass^5 42.1%Grok Voice 2 · Cost per hour $6.40/hr · pass^5 34.7% · † $0.08 per audio minute plus xAI's published $0.004 per text input, one per tool result. Without the text-input line Grok is $4.80/hr.Qwen Audio · Cost per hour $1.67/hr · pass^5 31.4%Nova Sonic 2 · Cost per hour $0.750/hr · pass^5 26.9%Flash Live 3.1 · Cost per hour $2.42/hr · pass^5 25.5% · † 22% of turns sent no usage packet; stage-final ones are billed from measured audio and re-sent context at Google's published rates (12% of the row), the rest are folded into the next packet by Gemini.Cascaded · Cost per hour $2.30/hr · pass^5 24.5%Realtime 2.1 Mini · Cost per hour $1.31/hr · pass^5 16.6%2.5 Flash Native · Cost per hour $1.06/hr · pass^5 12.0% · † 38% of turns sent no usage packet; those are billed from measured audio seconds (32 tokens/s), tool text and re-sent context at Google's published rates (33% of the row).GPT Live 1 (astra-med)GPT Live 1 (sol-low)Gemini 3.8 LiveQwen AudioNova Sonic 2Realtime 2.1Grok Voice 2 †Flash Live 3.1 †CascadedRealtime 2.1 Mini2.5 Flash Native †COST PER HOUR · USD / HR · ← betterPASS^5 · RELIABILITY, % · better →

Cost versus latency

4 of 12 harnesses sit on the frontier: Nova Sonic 2 ($0.750/hr, 2.54s), Realtime 2.1 Mini ($1.31/hr, 2.48s), Qwen Audio ($1.67/hr, 2.25s), and GPT Live 1 (sol-low) ($4.72/hr, 1.78s). Nova Sonic 2 (26.9%) and Realtime 2.1 Mini (16.6%) make it on price alone with pass5 under 30%.

Latency versus accuracy

Only GPT Live 1 (sol-low) (1.78s, 63.7%) and GPT Live 1 (astra-med) (1.81s, 66.3%) are on the frontier; everything else is slower and less reliable than at least one of them. Between them the trade is 37 ms of median latency for 2.6 points of pass5.

Cost versus accuracy

5 of 12 harnesses form the frontier: Nova Sonic 2 ($0.750/hr, 26.9%), Qwen Audio ($1.67/hr, 31.4%), Gemini 3.8 Live ($2.87/hr, 43.5%), GPT Live 1 (sol-low) ($4.72/hr, 63.7%), and GPT Live 1 (astra-med) ($6.46/hr, 66.3%). Best step: Gemini 3.8 Live → GPT Live 1 (sol-low), 20.2 points of pass5 for $1.85/hr more; priciest: GPT Live 1 (sol-low) → GPT Live 1 (astra-med), 2.6 points for $1.74/hr.

Rankings

One metric at a time

Best first, at cross-industry scope.

Latency

median s · p90 tick
0.0s1.0s2.0s3.0s4.0s5.0s6.0sGPT Live 1 (sol-low) · Median 1.78s · p90 2.11sGPT Live 1 (sol-low)1.78s2.11sGPT Live 1 (astra-med) · Median 1.81s · p90 2.19sGPT Live 1 (astra-med)1.81s2.19sRealtime 2.1 · Median 2.05s · p90 3.48sRealtime 2.12.05s3.48sQwen Audio · Median 2.25s · p90 3.31sQwen Audio2.25s3.31sFlash Live 3.1 · Median 2.36s · p90 3.67sFlash Live 3.12.36s3.67sRealtime 2.1 Mini · Median 2.48s · p90 4.03sRealtime 2.1 Mini2.48s4.03sGemini 3.8 Live · Median 2.48s · p90 3.22sGemini 3.8 Live2.48s3.22sGemini 3.8 Extended · Median 2.49s · p90 3.51sGemini 3.8 Extended2.49s3.51sNova Sonic 2 · Median 2.54s · p90 4.98sNova Sonic 22.54s4.98sGrok Voice 2 · Median 2.70s · p90 3.74sGrok Voice 22.70s3.74sCascaded · Median 3.10s · p90 4.95sCascaded3.10s4.95s2.5 Flash Native · Median 3.89s · p90 5.39s2.5 Flash Native3.89s5.39sMEDIAN LATENCY · VOICE-TO-VOICE · ← BETTER

Cost per hour

$/hr · list price
$0$1$2$3$4$5$6$7Nova Sonic 2 · Cost per hour $0.750/hrNova Sonic 2$0.750/hr2.5 Flash Native · Cost per hour $1.06/hr · † 38% of turns sent no usage packet; those are billed from measured audio seconds (32 tokens/s), tool text and re-sent context at Google's published rates (33% of the row).2.5 Flash Native †$1.06/hrRealtime 2.1 Mini · Cost per hour $1.31/hrRealtime 2.1 Mini$1.31/hrQwen Audio · Cost per hour $1.67/hrQwen Audio$1.67/hrCascaded · Cost per hour $2.30/hrCascaded$2.30/hrFlash Live 3.1 · Cost per hour $2.42/hr · † 22% of turns sent no usage packet; stage-final ones are billed from measured audio and re-sent context at Google's published rates (12% of the row), the rest are folded into the next packet by Gemini.Flash Live 3.1 †$2.42/hrGemini 3.8 Live · Cost per hour $2.87/hrGemini 3.8 Live$2.87/hrGPT Live 1 (sol-low) · Cost per hour $4.72/hrGPT Live 1 (sol-low)$4.72/hrRealtime 2.1 · Cost per hour $5.11/hrRealtime 2.1$5.11/hrGemini 3.8 Extended · Cost per hour $6.22/hr · † 32% of turns (tool calls) sent no usage packet; those are billed from measured audio seconds, tool text and re-sent context at Google's published rates (22.6% of the row). Thinking tokens on those turns are not recoverable.Gemini 3.8 Extended †$6.22/hrGrok Voice 2 · Cost per hour $6.40/hr · † $0.08 per audio minute plus xAI's published $0.004 per text input, one per tool result. Without the text-input line Grok is $4.80/hr.Grok Voice 2 †$6.40/hrGPT Live 1 (astra-med) · Cost per hour $6.46/hrGPT Live 1 (astra-med)$6.46/hrCOST PER HOUR · USD / HR · ← BETTER

pass5 by pack

  • Healthcare
  • Legal
  • Customer support
pass5 %
01020304050607080GPT Live 1 (astra-med) · Healthcare 77.5% · Legal 50.0% · Customer support 71.4%GPT Live 1 (astra-med)77.5%50.0%71.4%GPT Live 1 (sol-low) · Healthcare 78.3% · Legal 42.3% · Customer support 70.6%GPT Live 1 (sol-low)78.3%42.3%70.6%Gemini 3.8 Extended · Healthcare 63.9% · Legal 45.8% · Customer support 56.9%Gemini 3.8 Extended63.9%45.8%56.9%Gemini 3.8 Live · Healthcare 48.6% · Legal 25.0% · Customer support 56.9%Gemini 3.8 Live48.6%25.0%56.9%Realtime 2.1 · Healthcare 34.7% · Legal 50.0% · Customer support 41.7%Realtime 2.134.7%50.0%41.7%Grok Voice 2 · Healthcare 47.2% · Legal 18.1% · Customer support 38.9%Grok Voice 247.2%18.1%38.9%Qwen Audio · Healthcare 44.3% · Legal 16.7% · Customer support 33.3%Qwen Audio44.3%16.7%33.3%Nova Sonic 2 · Healthcare 36.1% · Legal 15.3% · Customer support 29.2%Nova Sonic 236.1%15.3%29.2%Flash Live 3.1 · Healthcare 34.7% · Legal 11.1% · Customer support 30.6%Flash Live 3.134.7%11.1%30.6%Cascaded · Healthcare 30.6% · Legal 9.7% · Customer support 33.3%Cascaded30.6%9.7%33.3%Realtime 2.1 Mini · Healthcare 19.1% · Legal 1.4% · Customer support 29.2%Realtime 2.1 Mini19.1%1.4%29.2%2.5 Flash Native · Healthcare 19.4% · Legal 0.0% · Customer support 16.7%2.5 Flash Native19.4%0.0%16.7%PASS^5 · RELIABILITY, % · BETTER →

Latency

GPT Live 1 (sol-low) is fastest at a 1.78s median; 2.5 Flash Native is slowest at 3.89s, 2.2× as long. Two harnesses answer under two seconds (GPT Live 1 (sol-low) and GPT Live 1 (astra-med)), 8 sit between two and three, and 2 take three or more.

Cost per hour

List price runs from $0.750/hr for Nova Sonic 2 to $6.46/hr for GPT Live 1 (astra-med), a 8.6× spread. The cheapest route to a pass5 above 60% is GPT Live 1 (sol-low) at $4.72/hr, above 50% is GPT Live 1 (sol-low) at $4.72/hr, and above 40% is Gemini 3.8 Live at $2.87/hr.

pass5 by pack

Averaged over the 12 harnesses, pass5 is 44.5% in healthcare, 42.4% in customer support, and 23.8% in legal. Pack leaders: GPT Live 1 (sol-low) in healthcare (78.3%), GPT Live 1 (astra-med) and Realtime 2.1 in legal (50.0%), and GPT Live 1 (astra-med) in customer support (71.4%).

Table

Every metric, one row per harness

Cross-industry · 12 harnesses

166.3%87.8%88.1%95.0%92.3%1.81s$6.46/hr
263.7%85.5%85.8%94.6%90.4%1.78s$4.72/hr
355.5%81.9%84.2%94.6%88.4%2.49s$6.22/hr
4
43.5%72.7%74.7%94.5%84.5%2.48s$2.87/hr
542.1%62.2%71.6%88.6%75.8%2.05s$5.11/hr
634.7%52.7%60.8%88.4%70.2%2.70s$6.40/hr
731.4%50.9%57.2%82.5%70.1%2.25s$1.67/hr
826.9%45.6%52.3%79.7%63.4%2.54s$0.750/hr
9
25.5%47.0%49.6%86.7%72.2%2.36s$2.42/hr
10
24.5%35.8%38.1%70.8%65.2%3.10s$2.30/hr
1116.6%36.3%42.2%71.5%56.8%2.48s$1.31/hr
1212.0%28.5%30.3%61.2%55.0%3.89s$1.06/hr

† Cost/hr figures marked on the charts carry a pricing caveat:

  • Gemini 3.8 Extended — 32% of turns (tool calls) sent no usage packet; those are billed from measured audio seconds, tool text and re-sent context at Google's published rates (22.6% of the row). Thinking tokens on those turns are not recoverable.
  • Grok Voice 2 — $0.08 per audio minute plus xAI's published $0.004 per text input, one per tool result. Without the text-input line Grok is $4.80/hr.
  • Flash Live 3.1 — 22% of turns sent no usage packet; stage-final ones are billed from measured audio and re-sent context at Google's published rates (12% of the row), the rest are folded into the next packet by Gemini.
  • 2.5 Flash Native — 38% of turns sent no usage packet; those are billed from measured audio seconds (32 tokens/s), tool text and re-sent context at Google's published rates (33% of the row).
  1. Healthcare, legal, and customer-support pass@1 are conversation-level combined pass among pass/fail only, from the scored August 2026 five-run suites.
  2. pass^5 is the unweighted mean, across tasks, of C(c,5)/C(n,5), the unbiased estimate that all five trials pass (tau-bench, Yao et al. 2024), on packs that published five independent runs. Healthcare, legal, and customer support all published five.
  3. Cross-industry pass@1 is the unweighted mean of the three launch packs.
  4. Cross-industry easy, medium, and hard are the unweighted mean of those difficulty pass@1 scores across the three launch packs.
  5. Cost/hr is the provider's published list price applied to what each scored conversation actually used, summed over the pack and divided by the pack's scored hours. Every rate was re-checked against the provider's own pricing page on 2026-09-29 and is recorded with its URL in mivas-bench voice-agent-harnesses/s2s-model-pricing.json: OpenAI gpt-realtime-2.1 ($4 text / $32 audio input, $0.40 cached, $24 text / $64 audio output per 1M tokens) and -mini ($0.60 / $10, $0.06 / $0.30 cached, $2.40 / $20); Gemini 3.8 Live and 3.1 Flash Live ($0.75 text / $3 audio in, $4.50 text / $12 audio out) and 2.5 Flash Native Audio ($0.50 / $3 in, $2 / $12 out); Amazon Nova Sonic 2 ($0.33 text / $3 audio in, $2.75 / $12 out, AWS Price List API us-east-1); Qwen audio-3.0-realtime-plus ($0.80 text / $6.40 audio in, $6.40 / $24 out, Alibaba Cloud international); Grok $0.08 per audio minute plus $0.004 per text input; gpt-live-1 $0.05 per session minute; gpt-4.1 $2 / $0.50 cached / $8; Deepgram Flux $0.0077 per minute (regular, not the promo); ElevenLabs Flash 2.5 $0.04 per 1,000 characters (API price list).
  6. Token models are priced from the usage packet of every traced model turn. A turn whose packet never arrived (the socket closed before turn_complete, or the plugin dropped it) is not left unbilled: its reply audio is billed at the provider's published audio price on the measured playout seconds (Gemini $0.018 per output minute and $0.005 per input minute; 32 tokens per second for 2.5, 20 and 10 tokens per second for OpenAI, 12.5 for Qwen), its tool-call text at four characters per token, and its context at the text lane of the neighbouring packet. Share of turns priced this way: Gemini 3.8 Live Extended Thinking 31.6% (8,223 of 26,026 turns, 22.6% of its cost), 2.5 Flash Native Audio 38.4% (3,901 of 10,151, 33.4%), Flash Live 3.1 21.6% (1,989 of 9,229, 11.6%; 3.1 folds a tool turn into the next packet, so only stage-final turns are added), Gemini 3.8 Live 1.0% (140 of 13,562), gpt-realtime-2.1-mini 0.4%, Qwen 0.2%. Nova, gpt-realtime-2.1 and the cascaded baseline needed none. Twenty conversations with no trace at all are priced from the exporter's token columns; thirteen with neither are rebuilt turn by turn from the transcript with the same published rates.
  7. Grok is metered at $0.08 per audio minute plus xAI's published $0.004 per text input; the harness sends one text input per tool result (13,774 across the three packs), which is the line that lifts Grok above $4.80/hr. Whether xAI counts function results as text inputs is not spelled out on its price sheet.
  8. GPT Live 1 (astra-medium) and GPT Live 1 (sol-low) are the September 2026 five-run suites on healthcare, legal, and customer support. Both use the gpt-live-1 voice layer; astra-medium delegates to gpt-6-astra at medium reasoning effort, sol-low to gpt-5.6-sol at low.
  9. GPT Live 1 cost/hr is the whole tested system at list price: $0.05 per minute for the gpt-live-1 voice session on the API-reported session seconds, plus the traced backend tokens (gpt-6-astra $10 input, $1 cached, $50 output per 1M; gpt-5.6-sol $4, $0.40, $20). Calls the voice layer handled without delegating carry no backend tokens and are billed for the session minutes only.
  10. Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are the September 2026 five-run suites on healthcare, legal, and customer support through the Gemini Live API; extended runs thinking_level HIGH. Their prompts end with the runtime's calendar block, added 2026-09-28 after plain 3.8 Live substituted the real-world date for the pack date once a session carried handoff history; the plain legal row is the rerun with that block, the other Gemini rows are the original runs.
  11. Gemini 3.8 cost/hr is list price on traced per-turn usage: $0.75 text and $3.00 audio per 1M input tokens, $4.50 text and $12.00 audio per 1M output tokens, output including thinking tokens. Usage is read from the realtime_inference span of each model turn and the child realtime_metrics span when the packet landed late. Turns with no packet (1% on 3.8 Live, 32% on Extended Thinking, nearly all tool-call turns) are billed from their measured audio seconds, tool-call text and re-sent context at the same published rates, so the figure is no longer a floor; the one thing it cannot recover is thinking tokens emitted on a packet-less turn. Gemini Flash Live 3.1 and 2.5 Flash Native Audio were recomputed the same way (2026-09-29), which also removed an earlier delta-decoding bug that had halved their per-turn context.
  12. Qwen audio-3.0-realtime-plus was previously priced at a proxy row (qwen3-omni-flash-realtime) with its whole input billed as audio; it now uses its own verified Alibaba row with the input split into caller audio (12.5 tokens per second, re-sent every turn per Alibaba's billing rules) and text, which is why its cost/hr fell about threefold.
  13. Cascaded Baseline is a cascaded STT/LLM/TTS stack: Deepgram Flux flux-general-en, OpenAI gpt-4.1, and ElevenLabs Flash 2.5 (eleven_flash_v2_5).