Bluejay / Labs
Open navigation

Working paper / 2026

MIVAS

Multi-Industry Voice Agent Simulation Bench

Abstract

Multi-Industry Voice Agent Simulation Bench (MIVAS) is an indicator of voice AI performance across economic sectors. MIVAS is comprised of tasks and verifiers with granular rewards executed in production-grade multi-agent RL environments across industries where voice AI is being adopted the fastest. Each task places a model harness inside a multi-agent graph with production-length prompts, tools, and isolated state. Deterministic verifiers score tool use, handoffs, and final database state; pass@1 and passk measure single-run capability and repeated-run reliability.

GitHub repository ↗

Clone the repository, pair a model harness with an industry task suite, and reproduce the evaluation environment.

Back to Model Harnesses

Cascaded Baseline

Cascaded STT / LLM / TTS · flux-general-en + gpt-4.1 + eleven_flash_v2_5

What the model is

Not a speech-to-speech model. The scored stack is Deepgram Flux flux-general-en for STT, OpenAI gpt-4.1 for the LLM, and ElevenLabs Flash 2.5 (eleven_flash_v2_5) for TTS. The same STT/LLM/TTS trio is what the Vapi and Cartesia cascaded harnesses use; here the framework is LiveKit Agents.

Harness in mivas-bench

How the harness was built

voice-agent-harnesses/livekit/cascaded/agent.py wires AgentSession(stt=deepgram.STTv2(model="flux-general-en"), llm=openai.LLM(model="gpt-4.1"), tts=elevenlabs.TTS(model="eleven_flash_v2_5"), vad=silero). Shared livekit/harness.py loads the industry blueprint and runs tools in-process.

The simulation engine dials SIP into LiveKit Cloud. An inbound trunk plus dispatch rule create the room and dispatch mivas-livekit-cascaded (or mivas-{slug} on Kubernetes). Audio is the stock LiveKit SIP mix. There is no CHIRP WebSocket.

Multi-agent assumptions

Handoff is in-framework: a handoff tool returns the target BlueprintAgent. History stays on the AgentSession. Industry tools POST to the state API from the worker process. Session tools such as end_call hang up after farewell playout.

Runtime

STT: Deepgram Flux flux-general-en. LLM: OpenAI gpt-4.1. TTS: ElevenLabs Flash 2.5, voice 21m00Tcm4TlvDq8ikWAM (Rachel). VAD: Silero. max_tool_steps=16.

One LiveKit worker process takes the SIP job. Kubernetes uses the LiveKit worker box (1000m / 1Gi / 3Gi). Tool POST timeout: 30 s.

Deployment

LiveKit worker Deployment, MIVAS_MODE=agent. run.py --harness livekit/cascaded --apply. The simulation runner uses connection_type=SIP. The harness needs DEEPGRAM_API_KEY, OPENAI_API_KEY, ELEVENLABS_API_KEY, and the LiveKit Cloud SIP secret.

Source: mivas-bench, run.py, k8s/deployment.yaml.

Scores

pass@1 among pass/fail only. pass5 is reliability from the five-run healthcare, legal, and customer-support suites.

Industrypass@1pass5ToolsHandoffFinal DBLatencyCost/hr
Healthcarescored38.9%30.6%38.9%73.6%73.3%2.51s$2.60/hr
Legalscored32.5%9.7%36.1%79.2%42.8%2.95s$2.22/hr
Customer supportscored36.1%33.3%39.4%59.7%79.4%3.85s$2.08/hr
Cross-industry35.8%24.5%38.1%70.8%65.2%3.10s$2.30/hr