Bluejay / Labs
Open navigation

Working paper / 2026

MIVAS

Multi-Industry Voice Agent Simulation Bench

Abstract

Multi-Industry Voice Agent Simulation Bench (MIVAS) is an indicator of voice AI performance across economic sectors. MIVAS is comprised of tasks and verifiers with granular rewards executed in production-grade multi-agent RL environments across industries where voice AI is being adopted the fastest. Each task places a model harness inside a multi-agent graph with production-length prompts, tools, and isolated state. Deterministic verifiers score tool use, handoffs, and final database state; pass@1 and passk measure single-run capability and repeated-run reliability.

GitHub repository ↗

Clone the repository, pair a model harness with an industry task suite, and reproduce the evaluation environment.

Back to Model Harnesses

Qwen Audio Realtime

Speech-to-speech · qwen-audio-3.0-realtime-plus

What the model is

Alibaba Cloud Model Studio / DashScope qwen-audio-3.0-realtime-plus (QWEN_AUDIO_MODEL). Frontend-only speech-to-speech; the harness does not attach an ACP coding-agent backend. The scored deploy uses a Singapore (ap-southeast-1) workspace. The model is offered in Beijing and Singapore only.

Provider model card · Harness in mivas-bench

How the harness was built

voice-agent-harnesses/qwen/ is a DashScope Realtime WebSocket client. The socket URL is built from QWEN_WORKSPACE_ID and QWEN_REGION, or from QWEN_WS_URL. Industry tools POST to the state API. Session tools hang up locally. Tools are declared in the nested Model Studio shape {type: function, function: {name, description, parameters}}.

Speak-first seeds a user conversation.item.create, then response.create. Qwen rejects a bare create on an empty conversation. Greeting text stays in the pack. Audio is Qwen-Audio PCM 16 kHz in / 24 kHz out, bridged to CHIRP 16 kHz pcm_s16le. Barge-in is Qwen server_vad.

Multi-agent assumptions

Handoffs are soft session.update on the same WebSocket, history stays. That matches the Grok pattern: the industry DAG swaps instructions and tools in place. Today's date is injected. infer_schedule_appointment recovers verbal booking confirms without a function-call event.

Runtime

server_vad defaults: threshold 0.5, silence_duration_ms 800 (QWEN_VAD_*). Session.update wait timeout: 60 s. Tool POST timeout: 30 s. Default voice longanqian, honored on the first session.update only. Local CHIRP port 8769 (8765 in-pod). Kubernetes resources are the default CHIRP box.

Deployment

CHIRP family on Kubernetes: run.py --harness qwen/audio-realtime --apply. Needs DASHSCOPE_API_KEY or QWEN_API_KEY, plus workspace id or QWEN_WS_URL, on mivas-secrets. The deployment assigns a dedicated WebSocket hostname to each harness and industry pair.

Source: mivas-bench, run.py, k8s/deployment.yaml.

Scores

pass@1 among pass/fail only. pass5 is reliability from the five-run healthcare, legal, and customer-support suites.

Industrypass@1pass5ToolsHandoffFinal DBLatencyCost/hr
Healthcarescored63.2%44.3%63.2%85.7%83.4%2.14s$1.82/hr
Legalscored51.7%16.7%54.7%83.9%58.3%2.38s$1.66/hr
Customer supportscored37.8%33.3%53.6%77.8%68.6%2.24s$1.54/hr
Cross-industry50.9%31.4%57.2%82.5%70.1%2.25s$1.67/hr