Public evaluation directory
Benchmarks
Reproducible evaluations built around the conditions AI systems encounter in production, complete with task catalogs, traces, scoring gates, and methodology.
Active programs / 01
MIVAS
Preview / LiveA suite of stateful voice-agent tasks spanning healthcare, legal, and customer support. Model harnesses operate inside multi-agent graphs and are scored with deterministic trajectory verifiers.
↗
Cross-industry leader
GPT Live 1 (astra-medium)
87.8%