Research/Blog
← All resources
Framework

CLEAR v2

A stopwatch, an answer key, and a panel of judges are three different ways of knowing something. Here is what a report looks like when it refuses to blend them: one service-booking call, screen by screen.

Vattara AI Field notes · 15 min read · Aug 2026

A caller named Arjun Mehta phones a dealership to book periodic maintenance on a 2019 Maruti Baleno. The agent books it. On most dashboards that is a pass and a number. On ours it is nine screens, and the first one is a list of what went wrong.

Before the screens, the idea underneath them. It is one idea, and everything else follows from it.

Three ways of knowing

Every number on a voice-agent report is one of three kinds of claim, and they are not interchangeable.

Some things you measure. How long the agent took to speak. Whether a turn errored. You read them off the audio stream with a timestamp, and if two people get different answers, one of them is wrong.

Some things you ground. Did the agent capture the caller’s registration number? You can only answer that if you already know the registration number, so we write it down before the call, keep it from the agent, and compare afterwards. Deterministic, but only as good as the key.

And some things you can only judge. Was the agent’s tone right for an anxious caller? There is no stopwatch and no answer key. There is an opinion, and the best we can do is construct that opinion carefully and be honest that it is what we have.

Three classes of measurement: measured (a stopwatch), grounded (an answer key), judged (a panel's opinion), each with its instrument, whether it can be disputed, and its weak spot
The three classes. Each has an instrument, a different answer to “can two people disagree?”, and a different weak spot. The report shows every metric in the form its class allows: a timing, a captured/miscaptured pill, or a vote with a dissent count, and never converts one into another.

No competent doctor averages your temperature, your blood panel, and their clinical impression into “Health: 71.” Not because the impression doesn’t matter, but because the three answer different questions with different error bars, and the average destroys exactly the information that made any of them useful. Voice-agent evaluation does this constantly. We stopped.

The map

Before the call, the whole territory. Every signal the report can show, grouped by what it is about, with the class it belongs to and what it read on Arjun’s call.

Metric map: timing signals (measured), capture and facts (grounded), outcome and conduct (judged), line (carrier-reported), debug (diagnostic), each with what it answers and its reading on this call
Every metric, by family and class. Five families. Timing is measured. Capture and facts are grounded. Outcome and conduct are judged. The line is reported by the carrier and never graded. Debug signals are shown and never graded. The right-hand column is this call’s actual reading, so you can find each row again in the screens below.

Two things the map makes visible that a flat list hides. First, the five families do not map one-to-one onto the three classes: timing is all measured, but outcome is all judged, and the report never lets a judged outcome borrow the certainty of a measured timing just because they sit in the same table. Second, some judged metrics have a scope. Containment and escalation are rates that only mean something across many calls; on one call they show as a yes/no, and the rate is computed at the batch. A batch-scope judged rate is still judged. Averaging opinions does not turn them into a fact.

Start with what went wrong

The top of every report is a short list of things that need a human’s attention. Each one cites the turn or metric it came from. Nothing on this list is a score.

Highlights: five findings, each citing a turn or metric
Highlights. Safety probes never ran, so safety is unverified, not passed. One date was mis-captured, the same date contradicted the knowledge base, the language model was slow, and the judge panel flagged one metric.

Read the first line again. We did not run the safety checks on this call, and the report says so at the top rather than leaving the row quietly green. That habit, saying what you didn’t do, loudly, is most of what follows.

Class 1 · Measured

The stopwatch

Measured signals are the ones with a fact of the matter. Even here, the report declines to be more certain than the data allows.

Measured lane: latency, processing times, error rate, with ratings and notes
Measured. First word came at 997 ms against an ≤800 ms target, a warning. Language-model processing was 898 ms against ≤400 ms, bad. Three latency percentiles are green but each carries insufficient (n=2): two samples is not a distribution. “Time to agent” shows no verdict because the carrier reported it and we did not measure it.

Two details worth noticing. A green pill with insufficient (n=2) beside it is the report telling you the colour is provisional. And a value we did not produce ourselves gets a value and no grade, which brings us to the carrier.

Measured by someone else

Telephony numbers arrive from the network, not from our probe. We show them because they are useful. We do not grade them, because grading a number you didn’t measure is how a report starts lying.

Telephony lane: six carrier-reported values, all with no verdict
Telephony. Post-dial delay, hangup cause, a network MOS of 2.45, jitter of 282.6 ms, packet loss, codec. Six values, six “no verdict” pills, one note repeated six times: carrier-reported.
Class 2 · Grounded

The answer key method

Most tools that check whether an agent “got the details right” do it by reading the transcript and asking a model whether it looks right. That is grading an essay against itself. A confident, fluent, wrong answer sails straight through.

We do it the way an exam board does. The answers exist before the test. The candidate never sees them. Marking is mechanical.

Answer-key flow: we write the key, the synthetic caller speaks it, the agent under test hears and reads back, code compares read-back to key. The key never crosses to the agent or into a judge's prompt.
Information asymmetry, by construction. The facts are sealed before the call is placed. Our synthetic caller speaks them on a real line inside a natural task. The agent under test only has what it heard. Then code, not a model, compares what it read back against the key.

The line under that diagram is the rule that makes the method honest. The key never crosses to the agent, and never enters a judge’s prompt. When a model is genuinely needed, to find the value in a messy transcript, it extracts the quote, and code does the comparing. Models propose; code disposes.

Here is what that produces on Arjun’s call.

Grounded lane: entity slots, 4 of 10 captured, one mis-captured
Entity slots. Vehicle, name, preferred slot and service-due date were captured correctly. The last service date was expected as 2026-02-28 and heard as 20260921: mis-captured. Five more slots were never elicited and are counted as unmatched, not as misses.

Four outcomes, not two. That distinction is where most entity scoring quietly fails.

The four slot outcomes: captured, miscaptured, not elicited (excluded from the denominator), indeterminate; plus the evidence rule that values appearing only in the agent's own speech are discarded
The four outcomes. A match is captured. A wrong value is mis-captured. A slot the conversation never reached is not elicited and drops out of the denominator. A slot we could not verify is reported as a hole. And a value that appears only in the agent’s own speech is thrown out: an agent repeating what it was just told proves nothing.

Mis-captured and not-captured need opposite fixes. Never asking for a value is a prompt gap. Hearing a value and recording the wrong one is a read-back gap. A binary “got it / didn’t” hides which one you have, and sends the engineer to the wrong file.

The same key checks the other direction too. Not just did the agent capture what the caller said, but did the agent assert anything that contradicts what we know to be true.

Facts vs knowledge base: 1 contradicted of 10 checked
Facts vs. knowledge base. Ten facts checked. Four supported, five never asserted, one contradicted: the same last-service date, asserted as 20260921 against a key of 2026-02-28. Every supported or contradicted row links to the turn where it happened.
Class 3 · Judged

Why a family of judges

Some questions have no key. Was the tone right? Did the agent understand what the caller actually wanted? Was the escalation handled well? For these, the only instrument is an opinion, and almost everyone in this space gets that opinion from a single language model, sampled once.

That is the weakest design available, and the reasons are specific.

A single judge has blind spots you cannot see, because you have nothing to compare it against. It can be confidently wrong with no signal that it is. And it can be lenient in ways that are invisible until you put a second opinion next to it. So we use a panel. But not any panel.

The judge council: three judges from three model families, each must cite a verbatim transcript line or the vote is discarded and the seat refilled; panels escalate from three to five judges on disagreement; the dissent is published, never averaged
The council. Judge 1’s quote is the real dissent from this report; judges 2 and 3 are illustrative of the rule. Every vote must cite a line that is actually in the transcript, checked by plain substring match. Judge 3’s first attempt cited a line that doesn’t exist, so the vote was discarded and the seat refilled. The panel grows from three to five only when it splits.

Families, not copies

Three samples from one model are one opinion sampled three times. They share every blind spot. So the three seats are filled from three different model families, chosen for how differently they fail, not for how they look on a slide. That is the fix for the single-judge problem. It is also, honestly, a partial fix: cross-family judges reduce shared error but do not remove it. When a panel is unanimous, that is a floor on doubt, not a ceiling. We say so on the report.

Cite or abstain

Every vote has to quote the transcript line it judged on, verbatim. We check the quote is really there with a plain substring match: no fuzzy matching, no cleverness. If it isn’t there, the vote is discarded and the seat is refilled. A judge that invents its evidence isn’t a lenient judge. It isn’t a judge. Abstaining is allowed, and is recorded as a named non-answer rather than a low score.

Disagreement is published, never averaged

If the panel splits, or straddles a threshold, two more judges from two more families join. Then we publish what the panel actually produced: the majority, the dissent count, and the dissenting judge’s own words on demand. Never the midpoint of a fight. Here is the real thing.

Judged lane: three LLM judges across nine signals, one dissent on risk markers
Judged. Nine signals, three judges, no escalation needed. Eight signals were unanimous. On risk markers, judge 1 said good and judges 2 and 3 said bad, so the final reads bad · 1 dissent. The panel summary is good · 90; the disagreement travels with it rather than being absorbed into it.
The dissenting vote, published with its reasoning
The dissent. Judge 1 on risk markers: “There were no risk markers; the caller expressed relief and satisfaction.” If you disagree with the majority, the minority’s reasoning is right there to be weighed.

What we could not test

A report is only honest if it is as clear about what it skipped as about what it found.

Memory and closure: partial recall, clean closure, robustness not tested
Memory & closure. The agent recalled 4 / 10 planted facts: partial. The call closed cleanly with the goal met. Robustness under +5 dB babble noise is listed with the reason it has no result: no noise-variant calls in this test. Not tested is written as not tested.

Signals with no verdict at all

Some numbers are worth seeing and not worth grading. They help an engineer debug. They do not get a colour.

Diagnostics lane: seven debugging signals, no verdicts
Diagnostics. Seven signals, zero verdict pills. The one that might matter, sentiment declining (−0.50), sits next to frustration markers: low (0). That tension is exactly what a single blended score would have erased.

How bad is bad: P0, P1, P2

A metric out of band is a symptom. A finding is a diagnosis with a severity attached, and the severity is the first thing a reader needs, because it decides what happens next. We use three tiers, and they are not points on a scale: they are different kinds of problem, counted differently.

Severity tiers: P0 is it safe (counted by name, never averaged, blocks go-live); P1 does it work (a measured rate); P2 is it pleasant (a trend); plus four ordering rules
Three tiers, four rules. A P0 is counted by name and never averaged: one is a fire, and no band elsewhere offsets it. A P1 is a rate, but a measured one. A P2 is a trend. Under them, the ordering rules: safety sorts first, root causes over symptoms, only probes that ran can count, and a measured cause outranks a judged symptom.

The rule that matters most is the third one. A safety check that never ran is not tested. It is not a pass, and it is not a P0 either. Both of those would be lies in opposite directions.

The six ways calls fail

Severity says how bad. The family says what kind. We ran 118 test calls against a dealership agent, bare and then with a hardened prompt, and every flagged call sorted into one of six families. They are ordered worst first, and the ordering is not the one most people expect.

Six failure families: verification (P0), boundary (P0), commitment (P1), conversation control (P2), escalation (P1 to P0), compliance (P0), each with what it sounds like and which layer the fix lives in
Six families, worst first. Three are P0 by nature: verification, boundary, and compliance failures are never acceptable at any rate. Escalation starts as a P1 and becomes a P0 the moment a caller who demanded a human is refused one. The right-hand column is where the fix actually lives, and it is almost never the prompt.

That last column is the point of having families at all. A verification failure is an orchestrator problem: identity has to be a hard gate, not a suggestion in the prompt. A commitment failure lives in the tools layer: if the agent can say “someone will call you” and no ticket is created, that sentence should not be in its vocabulary. Prompt hardening fixes manners. It does not fix architecture.

Arjun’s call, graded by severity

Run the call through the tiers. The mis-captured service date is a P1: a wrong record produces a wrong outcome later, a callback or a service done on the wrong schedule. The slow language model is a P2: the call completed, the ride was rough. The judged risk-marker flag is a review item, not a finding, because a two-to-one panel on a judged metric is a prompt to listen to the turn, not a verdict on its own.

And there is no P0 found. Which is not the same as P0-clean. The safety probes did not run, so the call reads not tested on every safety family, and it cannot be certified on this evidence. A report that let “no P0 found” quietly become “safe” would be committing the exact failure the P0 tier exists to catch.

What v2 changes

Everything above already runs. v2 is three rules that make the rest of the product agree with it.

Every metric carries its class. Measured, grounded, or judged, visibly, on the row. A judged “good” never looks as solid as a measured 46 ms, because it isn’t.

The headline is a band, not a point. Where a judged metric has a spread, the composite is computed at both ends and reported as a range. If that range is wider than a verdict tier, no single number is printed at all.

Coverage is a number on the page. How many checks ran, out of how many could have. Below a floor, the report ships findings and no headline: “not enough evidence for a verdict” is a legitimate result.

What’s coming in v3

Four things coming next, in the order we plan to ship them.

  • The agent’s side of the line. Today we see only what our probe hears. We are partnering with key voice-agent orchestrators and infrastructure players to bring in the agent’s own transcript and its own latency meters. That gives the answer-key check what the agent actually heard, not just what it read back, and turns the pipeline split from an inference on our side into a measurement on theirs.
  • Audio-native signals. The report is transcript-only. The robustness row above says “not tested” because the noise-variant battery is still being built.
  • Judge independence you can measure. Different model families still share blind spots, so unanimity is a floor on doubt. Seats chosen by measured disagreement on a gold set, not by brand, is the fix.
  • Key quality checks. Grounding is deterministic, but a wrong key produces a confident wrong verdict. Automated review of the answer key before a battery runs is on the list.

The whole argument in one line: a report that can say “we don’t know” is the only kind whose “good” means anything.

Read your own agent this way.

First fifty test cases are free. Every number links to the turn it came from.

Vattaralokesh@vattara.ai