Bluejay / Labs
Open navigation

Public evaluation directory

Benchmarks

Reproducible evaluations built around the conditions AI systems encounter in production, complete with task catalogs, traces, scoring gates, and methodology.

Active programs / 01

MIVAS

Preview / Live

A suite of stateful voice-agent tasks spanning healthcare, legal, and customer support. Model harnesses operate inside multi-agent graphs and are scored with deterministic trajectory verifiers.

↗

Cross-industry leader

GPT Live 1 (astra-medium)

87.8%