01
Safe
Evaluate systems where a mistake has consequences, not in a demo.
Bluejay Labs / Manifesto
Why we exist
Over the next decade, voice will become the default medium of interaction between humanity and technology. We believe speech is more intuitive than keyboards, touchscreens, controllers, or any other interface built to command a machine. As artificial systems surpass human intelligence, human control means more than anything. Our goal is to build training and evaluation instruments that encourage conversational intelligence models to prioritize human interests with superhuman intelligence.
Most labs are improving how models write code or do knowledge work: tasks a system can take minutes, even hours, to finish. Speech models have a constraint those systems do not: a latency budget. A real-time model has under three seconds to answer before the person on the line hangs up. Speech-to-speech models are categorically different from the cascaded pipeline, and they need to be evaluated that way if they are going to improve.
Our mission is to accelerate the adoption of safe, trustworthy, and reliable conversational intelligence. We will release the benchmarks and datasets that make trustworthy voice AI development public, mainstream, and easy to do well.
01
Evaluate systems where a mistake has consequences, not in a demo.
02
Publish methods, tasks, traces, and scores so claims can be inspected and reproduced.
03
Measure whether the same work is completed consistently, not only once.
First public instrument
MIVAS, the Multi-Industry Voice Agent Simulation Bench, is our first public benchmark. It tests whether a voice system is ready to be trusted with real work.