What the model is
xAI grok-voice-latest, the Voice Agent / Speech-to-Speech model. The alias resolves to grok-voice-think-fast-2.0. Transport is the xAI Realtime WebSocket at wss://api.x.ai/v1/realtime (GROK_WS_URL overrides the base).
How the harness was built
voice-agent-harnesses/grok/ is a custom WebSocket client, not the OpenAI Agents SDK. harness.py loads the blueprint, declares that agent's tools as xAI function tools, and drives the session with session.update / response.create / conversation.item.create.
CHIRP still presents 16 kHz pcm_s16le. The adapter resamples to Grok's 24 kHz PCM. After session.updated the harness sends a bare response.create. Greeting words come from the pack instructions. Industry tools POST to the state API. end_call is local plus the same 2.5 s delayed close.
CHIRP speech.started often fires on agent echo in the mixed recording path. Muting on that signal chops the greeting. The adapter treats inbound audio as echo while agent TTS is open, unless RMS clears GROK_USER_RMS_ON (default 350) for a real barge-over-TTS. Echo PCM is not forwarded; Grok server_vad would otherwise cancel the greeting.
Multi-agent assumptions
Grok has no native multi-agent handoff API. The harness keeps one WebSocket for the whole call. A handoff tool runs session.update with the target agent's instructions and tools only. Conversation history stays on the socket. Mid-call updates leave voice, VAD, and PCM format alone so the live socket is not reset.
That is the assumption the industry DAG rides on: Robin can transfer to scheduling or billing by swapping the prompt and the tool list. The pack still owns the prompts. Extra glue: if the model verbally confirms a booking without emitting a function-call event, infer_schedule_appointment recovers schedule_appointment arguments from the transcript (same recovery as Nova and Qwen).
Runtime
Audio: 24 kHz PCM in and out on Grok, 16 kHz on CHIRP. Default voice: eve. Input transcription: grok-transcribe, language_hint en. server_vad defaults: threshold 0.7, silence_duration_ms 700, prefix_padding_ms 400 (GROK_VAD_*). Echo suppress window: GROK_ECHO_SUPPRESS_S default 1.25 s. Session.update wait timeout: 60 s. Tool POST timeout: 30 s.
Today's date is appended to every agent's instructions. Local CHIRP port defaults to 8768. Kubernetes resources are the default CHIRP box (250m / 384Mi / 1536Mi).
Deployment
CHIRP family on Kubernetes: run.py --harness grok/voice --apply. Same Deployment + Service + ALB WebSocket as OpenAI. The pod needs GROK_API_KEY or XAI_API_KEY from mivas-secrets. Local: uv run python run.py --harness grok/voice --mode chirp.
Source: mivas-bench, run.py, k8s/deployment.yaml.