Bluejay / Labs
Open navigation

Working paper / 2026

MIVAS

Multi-Industry Voice Agent Simulation Bench

Abstract

Multi-Industry Voice Agent Simulation Bench (MIVAS) is an indicator of voice AI performance across economic sectors. MIVAS is comprised of tasks and verifiers with granular rewards executed in production-grade multi-agent RL environments across industries where voice AI is being adopted the fastest. Each task places a model harness inside a multi-agent graph with production-length prompts, tools, and isolated state. Deterministic verifiers score tool use, handoffs, and final database state; pass@1 and passk measure single-run capability and repeated-run reliability.

GitHub repository ↗

Clone the repository, pair a model harness with an industry task suite, and reproduce the evaluation environment.

Back to Model Harnesses

Amazon Nova Sonic 2

Speech-to-speech · amazon.nova-sonic-v2

What the model is

Amazon Nova 2 Sonic on Bedrock, model id amazon.nova-2-sonic-v1:0 (NOVA_SONIC_MODEL). The API is InvokeModelWithBidirectionalStream. Default region us-east-1 (NOVA_SONIC_REGION, kept separate from the snapshot bucket region).

Provider model card · Harness in mivas-bench

How the harness was built

voice-agent-harnesses/aws/harness.py talks Bedrock directly. Nova requires an open USER audio content stream for the whole prompt. While the caller is quiet the bridge feeds silent PCM (512 frames at 16 kHz, about 32 ms, NOVA_SONIC_SILENCE_CHUNK_S default 0.03 s) so the stream does not idle out. Silence alone yields usage events. Speak-first is an interactive USER text block (default ".") after that audio stream is live. The pack owns the greeting words.

Industry tools POST to the state API. Session tools hang up after the delayed close. Barge-in follows Nova's provider interrupted event. The adapter does not mute on CHIRP VAD alone.

Multi-agent assumptions

Nova fixes tools at promptStart. A specialist handoff therefore opens a new Bedrock stream for the target agent. The new stream would otherwise cold-open, so the harness seeds it with the caller's last line and a short continue-do-not-greet system note (handoff_seed_text). That is the glue the Robin / Hal / Kestrel DAG needs on Sonic.

infer_schedule_appointment is the same verbal-booking recovery used on Grok. Today's date is injected into every session's instructions. Default voice: matthew.

Runtime

Audio: Nova PCM 16 kHz in, 24 kHz out, bridged to CHIRP 16 kHz pcm_s16le. Tool POST timeout: 30 s. Hangup delay: 2.5 s. Local CHIRP port: 8774.

Amazon's Bedrock streaming library shares state across a whole process. Two calls in one interpreter cancel each other. adapters/chirp.py is a TCP parent that never imports that client; each accept spawns adapters/chirp_call.py in a fresh Python process and passes the socket fd. One container still takes many calls. CPU and memory bound the count.

Deployment

Same Kubernetes CHIRP Deployment as OpenAI and Grok: run.py --harness aws/nova-sonic-2 --apply. Credentials come from AWS_ACCESS_KEY_ID / SECRET / SESSION_TOKEN on mivas-secrets, or the pod's EKS Pod Identity / IRSA chain (ensure_aws_credentials copies the boto3 default chain into env for the experimental Bedrock SDK). Default resources 250m / 384Mi / 1536Mi. Process-per-call is inside that one replica.

Source: mivas-bench, run.py, k8s/deployment.yaml.

Scores

pass@1 among pass/fail only. pass5 is reliability from the five-run healthcare, legal, and customer-support suites.

Industrypass@1pass5ToolsHandoffFinal DBLatencyCost/hr
Healthcarescored48.9%36.1%50.0%78.9%73.1%2.38s$0.820/hr
Legalscored45.0%15.3%49.4%86.9%57.8%2.51s$0.730/hr
Customer supportscored42.8%29.2%57.5%73.3%59.4%2.74s$0.700/hr
Cross-industry45.6%26.9%52.3%79.7%63.4%2.54s$0.750/hr