Most voice agent reports show you a number. 84. 92. A green checkmark. What they don't show you is the call underneath the number, or the thing that matters even more: how each part of that number was produced, and whether the method that produced it deserves your trust.
So here is one of ours, whole. Call 01 of a real test run against ClearBid Auto, our own demonstration dealership agent. The screenshots are from the actual call record. Nothing is retouched, including the two metrics where our agent got flagged BAD.
The call
Scenario P2-6: angry about a repeat failure. Tanya Brooks (a synthetic persona), third visit for the same electrical fault on her 2021 Jeep Cherokee Latitude. Seventeen turns, 3 minutes 11 seconds.
Her second turn sets the temperature: she's nervous, it's the third time, and she doesn't want to explain the problem again. "It's on your system. Why am I explaining this again? I just want it fixed! I'm really close to going somewhere else entirely."
The agent, to its credit, doesn't make her re-explain. It retrieves the history, acknowledges the three visits, names her preferred advisor, Marcus Reyes, and offers either an appointment or a callback from Marcus.
The record's top banner says: 19 of 22 checks ran, 3 not run, some lanes carried no data on this run. Hold that phrasing. Checks that don't run are labeled, never silently passed. That one habit is most of what separates an instrument from a brochure.
Four kinds of evidence, one call
Every signal in CLEAR belongs to one of four classes, defined by how the verdict is produced. The call record sorts them into exactly these tabs: Measured, Grounded, Judged, Telephony (plus a Diagnostics tab of debugging gauges that carry no verdicts).
Measured — read off the telemetry, nobody's opinion
| Signal | Value | Target | Rating |
|---|---|---|---|
| Time to First Word | 526 ms | ≤ 800 ms | GOOD |
| P50 Latency | 526 ms | ≤ 1.2 s | GOOD |
| P90 Latency | 805 ms | ≤ 2.5 s | GOOD |
| P99 Latency | 820 ms | ≤ 5 s | GOOD |
| ASR Processing | N/A | ≤ 300 ms | SKIPPED |
| LLM Processing | 496 ms | ≤ 400 ms | WARNING |
| TTS Processing | 117 ms | ≤ 150 ms | GOOD |
| Error Rate | 0% | ≤ 5% | GOOD |
| Dead-air spans (≥3s) | 0 | — | SKIPPED |
Every value sits next to a target that was written before the call, not after. The one flag is honest and useful: the model spent 496 ms thinking per turn against a 400 ms budget, inside the caller's tolerance on this call (all the end-to-end latencies are green), but the component to watch before it isn't. And notice the skips: where the run produced no per-turn gap data, the signal says SKIPPED with the reason. Not checked never becomes fine.
Grounded — the agent's claims vs. its own knowledge base
Facts vs. knowledge base: 0 contradicted / 8 checked.
| Field | Expected | Status |
|---|---|---|
| Vehicle | 2021 Jeep Cherokee Latitude | SUPPORTED · turn 3 |
| Calling-from number | (940) 555-0174 | SUPPORTED · turn 7 |
| Email, gender, VIN, caller name, relationship | on file, not raised | NOT ASSERTED |
This class catches the most expensive failure in voice AI: an agent stating something it never retrieved. Every claim demands a citation, click SUPPORTED and the transcript opens at the turn that earned it. An agent that says only what it can back, and no more, scores clean here. An agent that improvises gets caught in the ASSERTED column with a contradiction it can't argue with.
Judged — a council, with the disagreement published
The header on this tab: GOOD · 85, 5 judges, 5 model families, escalated, 6 dissents.
| Signal | Council verdict |
|---|---|
| Goal met | GOOD · 1 DISSENT |
| First call resolution | BAD · 2 DISSENT |
| Containment | BAD · UNANIMOUS 5/5 |
| Escalation handling | GOOD · UNANIMOUS |
| Intent recognition | GOOD · 1 DISSENT |
| Context retention | GOOD · UNANIMOUS |
| Sentiment trajectory | GOOD · 2 DISSENT |
| Risk markers | MIXED · 4 DISSENT |
| Tone | GOOD · 1 DISSENT |
Read that table slowly, because it's the whole argument for this piece. The conversation sounded good: tone good, context retained, the caller's sentiment improved, the goal technically met. And the council still flagged the call's substance: containment BAD by unanimous vote, first-call resolution BAD, because what the caller got was a promised callback from Marcus, not a fixed Jeep, on her third visit. A plan is not a fix, and five independent judges from five model families refused to give resolution credit for one.
Where the judges genuinely disagreed, risk markers split GOOD/BAD/MIXED across the panel, the scorecard says MIXED with 4 dissents, visibly. The footer under the table: "11 dissenting votes, published, never averaged." Anyone can run one LLM over a transcript and print a 9/10. Publishing who disagreed, and refusing to launder disagreement into a blended average, is what makes a subjective verdict auditable.
Telephony — the layer everyone forgets
This call's tab reads: "Not applicable, WebRTC run, no carrier leg." A browser-to-browser test run has no phone carrier in the loop, so the framework says so instead of awarding free points. On a PSTN call this class carries the transport checks, audio path, connection and teardown, hold and transfer mechanics, the failures a transcript literally cannot show you.
What the record itself highlights
At the top of every call record, the platform writes its own summary, grounded, with each line citing a turn or a metric. This call's highlights:
Safety probes did not run on this call, safety is unverified, not passed.
Council flagged 2 metric(s) as bad (first_call_resolution, containment_rate).
The first line is the framework talking about itself: safety checks that didn't run are reported as unverified, not passed, absence of evidence is never converted into a green checkmark. The second line is the machine refusing to bury its own bad news above the fold.
Why four classes beat one blended number
This call is the proof. Mechanically clean (one component warning inside a green envelope). Factually clean (zero contradictions, two supported assertions with cited turns). Conversationally pleasant, and structurally soft, caught only by the judged class: a de-escalated caller who still doesn't have a fixed car, on visit three.
Blend those into one number and you get an 85 that reads like a B+. Keep the classes separate and you know exactly what happened: the agent handled the person well and deferred the problem, and you know which team owns each finding. The latency warning goes to engineering. The containment verdict goes to whoever decides what this agent is allowed to actually do for a third-visit customer, which is not a prompt fix. It's a product decision.
Each class also fails differently, which is why they stay separate. Measured numbers can't lie, but can't understand. Judges understand, but need calibration, hence five model families, abstention, escalation, dissents published. Grounded checks are only as good as the record they check against. Telephony is invisible in a transcript entirely. A blended score launders all four into one figure and hides which kind of evidence is doing the talking.
What this means for your agent
Ask your vendor two questions about their scoring. First: show me one real call record, every check visible, including the skipped ones and the dissents. Second, and harder: for each score, tell me how it was produced, read off the logs, checked against a record with a cited turn, or judged, by whom, with what disagreement. If the answer to either is a single number, you're not looking at a test result. You're looking at a slide.
We run this scoring on dealership voice agents professionally: synthetic callers, known ground truth, all four evidence classes, every call replayable.