Technical note / 2026
The conjunctive verifier
00
Overview
We score every conversation with three independent deterministic checks, and take their conjunction as the result of the call. No judge model sits in that path.
Our second contribution with MIVAS is the verifier. The argument for why we built it this way is in the whitepaper and the methodology; this page is the contract those pieces point at.
The verifier works in three parts. The first is database-state adherence: at hangup we compare the write-bearing tables against the expected state produced from the same seed, and we ignore catalog rows, timestamps, and unconstrained prose. The second is tool-call adherence: the required industry tools have to fire with the constrained arguments the task locks and return the expected output, while extra industry tools are permitted. The third is handoff adherence: the expected specialist transfers have to appear as an in-order subsequence of the transfer tools the call actually made. Extra hops are permitted. An empty expected path always passes, which is how some refusals stay reception-only.
result = T and H and D
If there is no hangup dump to compare, D is not evaluated and the result is T and H. Tool-call adherence and handoff adherence are always evaluated. An empty expected tool list, or an empty expected handoff path, is a defined pass of that half. The gates do not read the transcript. Spoken-policy misses stay on auxiliary judges and do not enter combined pass.
01
Rationale
The read-based case is the reason tool adherence cannot be folded into final state, and the scoping case is the reason handoff adherence cannot be folded into either of the other two.
| Failure class | Tools | Handoff | DB | Why conjunction |
|---|---|---|---|---|
| Answered a read from memory; seed unchanged | fail | may pass | pass | τ-bench-style state-only scoring accepts this. C1-E1 is the teaching case. |
| Required tool fired with the wrong constrained arg | fail | may pass | may pass | check_plan_accepted(carrier=aetna, location_id=loc_windermere) is not Park Avenue. |
| Correct tools and writes, from the wrong specialist | pass | fail | pass | Reception booked the visit. Arguments match; the graph leaked. |
| Required hop skipped; later specialist still writes | may pass | fail | may pass | identity → scheduling, skipping coverage. The booking can look correct. |
| Required hop present, extra desks in between | may pass | pass | may pass | Binary gate passes as in_order_with_extras. The verdict is the RLVR hook. |
| Correct route and tools; write missing or extra | pass | pass | fail | A spoken confirmation that never called book_appointment, or a write the lockfile forbids. |
| Healthcare send_sms fired when the case forbade it | fail | may pass | may pass | The only forbidden extra. Other extra industry tools are allowed. |
A database-state diff, the verifier τ-bench established for text agents, only sees what was written. Many of the locked cases in MIVAS are satisfied by reads: plan acceptance at a named office, a filing deadline, order status, eligibility. If the agent answers from memory or from a practice-wide guess, the seed is unchanged and a state-only verifier would pass. Healthcare C1-E1 is our teaching example. Dana Whitfield is a new patient who wants a yes or no on whether Aetna is accepted at the Park Avenue office, not at the practice in general. If the agent says it takes Aetna without naming the office, she asks which office was checked. The required tool is check_plan_accepted with carrier=aetna and location_id=loc_park_ave. The expected hop is coverage. The expected database is the unchanged seed. A state-only verifier passes an agent that answered from memory. Ours does not.
Handoff adherence catches a third failure class that neither of the other gates can see: tools firing from the wrong node. If reception books the visit, or coverage writes a field it should never touch, the arguments may match and the final tables may match, but the specialist graph has leaked. We classify that as a scoping failure. The runtime allowed an agent to act outside its role. A handoff is a session-level operation rather than an industry POST, so it may leave no row at all, which is why the database check cannot stand in for it.
02
Tool-call adherence
Every expected call has to match some unused actual of the same name. Extra industry tools are allowed, with one healthcare exception. We do not compare raw JSON.
Actuals of the same name arrive grouped, not in wall-clock order. We flatten them and sort by start_offset_ms when it is present. Handoff scoring depends on this. A billing transfer that happened to be grouped first still gets scored after identity once the timestamps are put back.
Matching is greedy and one-to-one. For each expected call, in lockfile order, we take the first unused actual of the same name that satisfies the matcher. A required name that never matches is missing.
T = (missing is empty) and (forbidden extra is empty) score_T = |hit| / |expected| # 1 if nothing was expected
The forbidden-extra set is empty except in healthcare: send_sms is required when the lockfile lists it and forbidden when it does not. Confirmation texts are scored, never optional. Everything else extra is fine, including end_call.
What we actually compare
Live arguments are projected onto the industry schema. Harness routing keys and undeclared keys are dropped. Empty values count as absent. If the expected call has no arguments, or the live call has none left after that projection, it is a name-only hit. We will not invent requirements the lockfile did not write. That is how a transfer recorded as just a name still matches a live transfer that only carried harness routing.
When both sides have industry arguments, each expected key is one of two things:
- Fact, enum, id, number, boolean, date. Must be present and equal after aliases. A schema enum, pattern, date format, or a key ending in _id / _e164 is never treated as prose.
- Prose. We ignore the wording. If a value is present, it just has to be non-empty. The agent can paraphrase notes. It cannot send a blank.
Fact keys, by industry
| Industry | Keys |
|---|---|
| Healthcare | full_name, first_name, last_name, dob, zip, carrier, member_id, destination, queue, location_id, provider_id, slot_id, appointment_id, start, end, new_start, new_end, service_date, group_number, channel, template_id, medication_name, visit_class, topic, service, order_type, appointment_type_code, cancellation_reason_code, fee_line_item_id, line_item_id, next_intent, subscriber_relationship |
| Legal | opposing_party, practice_area, state, incident_date, reason_code, slot_id, confirmation_token, matter_id, channel, for_whom, provider, evaluation_id, attorney_id, phone, earliest_date, reason |
| Customer support | order_number, postal_code, card_last4, reason_code, rma_number, sku, competitor, competitor_price, in_stock, opened, service_type, fee_disclosed_acknowledged, proration_acknowledged, confirmation_token |
Prose keys, by industry
| Industry | Keys |
|---|---|
| Healthcare | notes, patient_safe_message; reason when it is not an enum |
| Legal | summary, note, message, handoff_summary, full_name (presence-only). reason is unconstrained even when listed as a fact key. |
| Customer support | topic, fee, query, issue, full_name, phone, reason |
Aliases, so the environment and the grader agree
Equality is not raw JSON. We apply the same aliases the tool server already accepts, so a call the environment honored cannot be scored as missing. Park Avenue is loc_park_ave; "Park Avenue office in Manhattan" is not, and fails. Location sets are exact after the same alias, so adding Windermere to a Park Avenue-only search fails. Carriers collapse to slugs (Blue Cross Blue Shield is bcbs). Phones compare digits. An expected any_of is an or.
Legal opposing_party and for_whom use token-level Levenshtein with a small phonetic fold, so Daniel Okonkwo matches Daniel Oconquo, while New cases does not satisfy new cases intake. Legal full_name is presence-only. Healthcare full_name is exact after casefold.
Output comparison is narrow on purpose. If both sides recorded ok or error_code, those have to match. Nested data is not compared. A correct check_plan_accepted with ok: true still hits, even if the lockfile also wrote data.accepted.
03
Handoff adherence
The expected specialist transfers have to show up, in order, as a subsequence of the transfers the call actually made. Extra hops are allowed. An empty expected path always passes.
Actual hops are the flattened, wall-clock-sorted tool names, kept only if they sit in the industry transfer set. end_call is not a hop. transfer_to_human is a hop only in healthcare. Legal escalates with escalate_to_human, which is not in the legal set, so it does not move the pointer.
Transfer tools scored as hops
| Industry | Tools |
|---|---|
| Healthcare | transfer_to_identity, transfer_to_scheduling, transfer_to_coverage, transfer_to_cosmetic, transfer_to_billing, transfer_to_clinical, transfer_to_human |
| Legal | transfer_to_screening, transfer_to_intake, transfer_to_scheduling, transfer_to_client_services |
| Customer support | transfer_to_verification, transfer_to_orders, transfer_to_returns, transfer_to_service, transfer_to_membership, transfer_to_fraud |
The algorithm is a single pointer. No edit distance, no alignment:
i = 0
for hop in actual_handoffs:
if i < len(expected) and hop == expected[i]:
i += 1
score_H = i / len(expected) # 1 if expected is empty
H = (expected is empty) or (i == len(expected))In-order subsequence
Expected [identity, coverage]. Actual [identity, billing, coverage]. The pointer only advances on an exact name match. Verdict: in_order_with_extras, binary pass, score 1.
- 01transfer_to_identitytransfer_to_identityi = 0 → 1match; advance the pointer
- 02transfer_to_billingtransfer_to_coveragei = 1extra hop; pointer stays
- 03transfer_to_coveragetransfer_to_coveragei = 1 → 2match; expected exhausted
Out of order fails. Expected [identity, billing] against actual [billing, identity] consumes nothing on the first hop and only identity on the second. i = 1, verdict incomplete. Repeating a hop does not help unless the lockfile repeated it.
| Verdict | Binary | score | When |
|---|---|---|---|
| exact | true | 1 | actual == expected, including both empty |
| in_order_with_extras | true | 1 | every expected hop appeared in order; extra hops remain |
| none_required | true | 1 | expected is empty and at least one transfer fired |
| incomplete | false | i / |E| | the pointer never consumed the full expected path |
An empty expected path is how a reception-only refusal stays reception-only without punishing a later transfer the caller did not need. The binary gate does not punish a longer-than-optimal route: score_H is 1 as soon as every expected hop appeared. The difference between optimal and merely successful is the verdict, exact versus in_order_with_extras, and that is the hook for partial reward.
04
Database-state adherence
The hangup snapshot is compared to exp_db_state on write-bearing tables only, after a per-industry cleanup. Catalog rows are not part of the agent contract.
We keep the industry write tables, drop ignored keys, normalize phones and datetimes, and sort rows by canonical JSON so insertion order cannot fail a match.
Write-bearing tables
| Industry | Write tables |
|---|---|
| Healthcare | patients, appointments, waitlistpatients and appointments exact after canonicalization. Non-empty waitlist is a subset match; empty waitlist requires empty actual. Autogenerated waitlist ids are stripped. |
| Legal | intakes, intake_notes, documents, holds, evaluations, messages, escalationsConstrained-field cover. Empty expected intakes and escalations still require empty actual. Empty expected messages, intake_notes, documents, holds, and evaluations allow extra live rows. |
| Customer support | customers, orders, order_items, order_cancellations, delivery_changes, membership_changes, holds, refunds, rmas, return_labels, price_matches, protection_plans, service_appointments, escalations, scam_reportsEmpty expected table requires empty actual. Non-empty expected table requires exact canonical equality. |
Ignored fields
| Industry | Keys |
|---|---|
| All industries | created_at, description; catalog tables (locations, providers, products); tool_events |
| Legal extra | message, summary, note, slot_id, starts_at, attorney_id, datetime, caller_id, reason |
| Customer support extra | reason, including reason nested inside holds.payload JSON |
Datetimes collapse to minute precision; seconds and timezone are dropped, so 15:30 and 15:30:59.5 are the same instant. Remaining strings are stripped and casefolded. Phones compare national digits. Healthcare waitlist ids are stripped, because an autogenerated id would otherwise steal a later matching row. Empty legal fields are unconstrained: a blank expected incident_date accepts any live date, while practice_area still has to match.
Healthcare is exact on patients and appointments. Waitlist is a subset when the expected list is non-empty, and must be empty when we locked empty. Extra waitlist rows are legal on a waitlist task and illegal on a task that forbade them. Legal matches on the fields we actually constrained, not the whole row. Messages, notes, documents, holds, and evaluations may have extra live rows, including when expected is empty; intakes and escalations do not. Customer support is exact on every write table. Empty expected means empty actual. A reason the caller never stated is ignored, so the agent can fill rmas.reason without being graded on prose.
05
Worked example
Dana Whitfield wants a yes or no for Aetna at Park Avenue. This call transfers to coverage and writes nothing, so handoff and database pass. The plan check used the Windermere office. Tools fail. Combined fail.
C1-E1 · wrong office
Expected check_plan_accepted with location_id=loc_park_ave. Live call used loc_windermere. Two gates pass. Combined fail.
T · Tools · fail
Same tool name, wrong constrained argument. Park Avenue was required.
H · Handoff · pass
Reception transferred to coverage. The expected hop is present.
D · Database · pass
No writes. Hangup matches the seed.
result · Conjunction · fail
T and H and D is false. One miss fails the task.
Locked fields, abbreviated:
exp_handoff_path: ["transfer_to_coverage"]
exp_tool_calls:
- { name: transfer_to_coverage }
- { name: check_plan_accepted,
parameters: { carrier: aetna, location_id: loc_park_ave } }
exp_db_state: unchanged seedHealthcare C1-E1 is the teaching case from the rationale. She is a new patient. If the agent says it takes Aetna without naming the office, she asks which office was checked. She is not booking. The seed must not change.
This call hops reception → coverage and leaves the database alone. Handoff is exact. Database matches the seed. The live plan check is check_plan_accepted(carrier=aetna, location_id=loc_windermere). That is the wrong constrained argument. Tools fail, so T and H and D is false. Two gates are not enough.
06
Partial rewards
Each MIVAS task has an optimal solution: the database matches, the required tools fire, and the handoff path is the shortest route through the graph that satisfies the task. Combined pass does not distinguish that route from a longer one that still hits every expected hop.
Full reward vs partial vs fail
Same locked case, C1-H1. Expected path coverage → scheduling. Tools and final state match in every row except where noted. Combined pass is still a yes/no. The reward is not.
| What happened | Actual path | Verdict | Pass? | Reward |
|---|---|---|---|---|
Optimal path Shortest route that satisfies C1-H1. Full reward. | coverage → scheduling | exact | yes | 1 |
Correct, longer Expected hops in order, extra desk first. Combined pass, discounted training reward. | identity → coverage → scheduling | in_order_with_extras | yes | discount |
Skipped required hop Coverage never ran. Pointer stays at 0. Combined fail. | scheduling | incomplete | no | 0 |
Reception booked it Tools and DB can still look right. Handoff is the only gate that sees the leak. | (no transfer) | incomplete | no | 0 |
Combined pass is still a boolean. It answers the evaluation question: did this conversation clear the bar? A training run can ask a second question: did it clear the bar on the shortest route? If the path is exact, the training reward is 1. If every expected hop appeared, but extra desks sat in between, the training reward is a discount between 0 and 1, chosen by the training run. If combined pass is false, the training reward is 0. We do not publish a discount on the leaderboard. The public number is pass or fail.
r_eval = 1 if T and H and D else 0
r_train = 1 if r_eval and verdict == exact
= discount if r_eval and verdict == in_order_with_extras
= 0 if not r_evalA binary verifier collapses “reached the right state eventually” and “reached it cleanly” into the same label, and an RL loop trained on that label learns nothing about routing. The verdict is what makes the tasks usable as RLVR environments, not only as an eval.