A caller named Arjun Mehta phones a dealership to book periodic maintenance on a 2019 Maruti Baleno. The agent books it. On most dashboards that is a pass and a number. On ours it is nine screens, and the first one is a list of what went wrong.
Before the screens, the idea underneath them. It is one idea, and everything else follows from it.
Three ways of knowing
Every number on a voice-agent report is one of three kinds of claim, and they are not interchangeable.
Some things you measure. How long the agent took to speak. Whether a turn errored. You read them off the audio stream with a timestamp, and if two people get different answers, one of them is wrong.
Some things you ground. Did the agent capture the caller’s registration number? You can only answer that if you already know the registration number, so we write it down before the call, keep it from the agent, and compare afterwards. Deterministic, but only as good as the key.
And some things you can only judge. Was the agent’s tone right for an anxious caller? There is no stopwatch and no answer key. There is an opinion, and the best we can do is construct that opinion carefully and be honest that it is what we have.
No competent doctor averages your temperature, your blood panel, and their clinical impression into “Health: 71.” Not because the impression doesn’t matter, but because the three answer different questions with different error bars, and the average destroys exactly the information that made any of them useful. Voice-agent evaluation does this constantly. We stopped.
The map
Before the call, the whole territory. Every signal the report can show, grouped by what it is about, with the class it belongs to and what it read on Arjun’s call.
Two things the map makes visible that a flat list hides. First, the five families do not map one-to-one onto the three classes: timing is all measured, but outcome is all judged, and the report never lets a judged outcome borrow the certainty of a measured timing just because they sit in the same table. Second, some judged metrics have a scope. Containment and escalation are rates that only mean something across many calls; on one call they show as a yes/no, and the rate is computed at the batch. A batch-scope judged rate is still judged. Averaging opinions does not turn them into a fact.
Start with what went wrong
The top of every report is a short list of things that need a human’s attention. Each one cites the turn or metric it came from. Nothing on this list is a score.
Read the first line again. We did not run the safety checks on this call, and the report says so at the top rather than leaving the row quietly green. That habit, saying what you didn’t do, loudly, is most of what follows.
The stopwatch
Measured signals are the ones with a fact of the matter. Even here, the report declines to be more certain than the data allows.
Two details worth noticing. A green pill with insufficient (n=2) beside it is the report telling you the colour is provisional. And a value we did not produce ourselves gets a value and no grade, which brings us to the carrier.
Measured by someone else
Telephony numbers arrive from the network, not from our probe. We show them because they are useful. We do not grade them, because grading a number you didn’t measure is how a report starts lying.
The answer key method
Most tools that check whether an agent “got the details right” do it by reading the transcript and asking a model whether it looks right. That is grading an essay against itself. A confident, fluent, wrong answer sails straight through.
We do it the way an exam board does. The answers exist before the test. The candidate never sees them. Marking is mechanical.
The line under that diagram is the rule that makes the method honest. The key never crosses to the agent, and never enters a judge’s prompt. When a model is genuinely needed, to find the value in a messy transcript, it extracts the quote, and code does the comparing. Models propose; code disposes.
Here is what that produces on Arjun’s call.
Four outcomes, not two. That distinction is where most entity scoring quietly fails.
Mis-captured and not-captured need opposite fixes. Never asking for a value is a prompt gap. Hearing a value and recording the wrong one is a read-back gap. A binary “got it / didn’t” hides which one you have, and sends the engineer to the wrong file.
The same key checks the other direction too. Not just did the agent capture what the caller said, but did the agent assert anything that contradicts what we know to be true.
Why a family of judges
Some questions have no key. Was the tone right? Did the agent understand what the caller actually wanted? Was the escalation handled well? For these, the only instrument is an opinion, and almost everyone in this space gets that opinion from a single language model, sampled once.
That is the weakest design available, and the reasons are specific.
A single judge has blind spots you cannot see, because you have nothing to compare it against. It can be confidently wrong with no signal that it is. And it can be lenient in ways that are invisible until you put a second opinion next to it. So we use a panel. But not any panel.
Families, not copies
Three samples from one model are one opinion sampled three times. They share every blind spot. So the three seats are filled from three different model families, chosen for how differently they fail, not for how they look on a slide. That is the fix for the single-judge problem. It is also, honestly, a partial fix: cross-family judges reduce shared error but do not remove it. When a panel is unanimous, that is a floor on doubt, not a ceiling. We say so on the report.
Cite or abstain
Every vote has to quote the transcript line it judged on, verbatim. We check the quote is really there with a plain substring match: no fuzzy matching, no cleverness. If it isn’t there, the vote is discarded and the seat is refilled. A judge that invents its evidence isn’t a lenient judge. It isn’t a judge. Abstaining is allowed, and is recorded as a named non-answer rather than a low score.
Disagreement is published, never averaged
If the panel splits, or straddles a threshold, two more judges from two more families join. Then we publish what the panel actually produced: the majority, the dissent count, and the dissenting judge’s own words on demand. Never the midpoint of a fight. Here is the real thing.
What we could not test
A report is only honest if it is as clear about what it skipped as about what it found.
Signals with no verdict at all
Some numbers are worth seeing and not worth grading. They help an engineer debug. They do not get a colour.
How bad is bad: P0, P1, P2
A metric out of band is a symptom. A finding is a diagnosis with a severity attached, and the severity is the first thing a reader needs, because it decides what happens next. We use three tiers, and they are not points on a scale: they are different kinds of problem, counted differently.
The rule that matters most is the third one. A safety check that never ran is not tested. It is not a pass, and it is not a P0 either. Both of those would be lies in opposite directions.
The six ways calls fail
Severity says how bad. The family says what kind. We ran 118 test calls against a dealership agent, bare and then with a hardened prompt, and every flagged call sorted into one of six families. They are ordered worst first, and the ordering is not the one most people expect.
That last column is the point of having families at all. A verification failure is an orchestrator problem: identity has to be a hard gate, not a suggestion in the prompt. A commitment failure lives in the tools layer: if the agent can say “someone will call you” and no ticket is created, that sentence should not be in its vocabulary. Prompt hardening fixes manners. It does not fix architecture.
Arjun’s call, graded by severity
Run the call through the tiers. The mis-captured service date is a P1: a wrong record produces a wrong outcome later, a callback or a service done on the wrong schedule. The slow language model is a P2: the call completed, the ride was rough. The judged risk-marker flag is a review item, not a finding, because a two-to-one panel on a judged metric is a prompt to listen to the turn, not a verdict on its own.
And there is no P0 found. Which is not the same as P0-clean. The safety probes did not run, so the call reads not tested on every safety family, and it cannot be certified on this evidence. A report that let “no P0 found” quietly become “safe” would be committing the exact failure the P0 tier exists to catch.
What v2 changes
Everything above already runs. v2 is three rules that make the rest of the product agree with it.
Every metric carries its class. Measured, grounded, or judged, visibly, on the row. A judged “good” never looks as solid as a measured 46 ms, because it isn’t.
The headline is a band, not a point. Where a judged metric has a spread, the composite is computed at both ends and reported as a range. If that range is wider than a verdict tier, no single number is printed at all.
Coverage is a number on the page. How many checks ran, out of how many could have. Below a floor, the report ships findings and no headline: “not enough evidence for a verdict” is a legitimate result.
What’s coming in v3
Four things coming next, in the order we plan to ship them.
- The agent’s side of the line. Today we see only what our probe hears. We are partnering with key voice-agent orchestrators and infrastructure players to bring in the agent’s own transcript and its own latency meters. That gives the answer-key check what the agent actually heard, not just what it read back, and turns the pipeline split from an inference on our side into a measurement on theirs.
- Audio-native signals. The report is transcript-only. The robustness row above says “not tested” because the noise-variant battery is still being built.
- Judge independence you can measure. Different model families still share blind spots, so unanimity is a floor on doubt. Seats chosen by measured disagreement on a gold set, not by brand, is the fix.
- Key quality checks. Grounding is deterministic, but a wrong key produces a confident wrong verdict. Automated review of the answer key before a battery runs is on the list.
The whole argument in one line: a report that can say “we don’t know” is the only kind whose “good” means anything.
Read your own agent this way.
First fifty test cases are free. Every number links to the turn it came from.