What the model is
Not a speech-to-speech model. The scored stack is Deepgram Flux flux-general-en for STT, OpenAI gpt-4.1 for the LLM, and ElevenLabs Flash 2.5 (eleven_flash_v2_5) for TTS. The same STT/LLM/TTS trio is what the Vapi and Cartesia cascaded harnesses use; here the framework is LiveKit Agents.
How the harness was built
voice-agent-harnesses/livekit/cascaded/agent.py wires AgentSession(stt=deepgram.STTv2(model="flux-general-en"), llm=openai.LLM(model="gpt-4.1"), tts=elevenlabs.TTS(model="eleven_flash_v2_5"), vad=silero). Shared livekit/harness.py loads the industry blueprint and runs tools in-process.
The simulation engine dials SIP into LiveKit Cloud. An inbound trunk plus dispatch rule create the room and dispatch mivas-livekit-cascaded (or mivas-{slug} on Kubernetes). Audio is the stock LiveKit SIP mix. There is no CHIRP WebSocket.
Multi-agent assumptions
Handoff is in-framework: a handoff tool returns the target BlueprintAgent. History stays on the AgentSession. Industry tools POST to the state API from the worker process. Session tools such as end_call hang up after farewell playout.
Runtime
STT: Deepgram Flux flux-general-en. LLM: OpenAI gpt-4.1. TTS: ElevenLabs Flash 2.5, voice 21m00Tcm4TlvDq8ikWAM (Rachel). VAD: Silero. max_tool_steps=16.
One LiveKit worker process takes the SIP job. Kubernetes uses the LiveKit worker box (1000m / 1Gi / 3Gi). Tool POST timeout: 30 s.
Deployment
LiveKit worker Deployment, MIVAS_MODE=agent. run.py --harness livekit/cascaded --apply. The simulation runner uses connection_type=SIP. The harness needs DEEPGRAM_API_KEY, OPENAI_API_KEY, ELEVENLABS_API_KEY, and the LiveKit Cloud SIP secret.
Source: mivas-bench, run.py, k8s/deployment.yaml.