Research/Blog
← All resources
Blog

We Scored One of Our Own Calls. Here's the Actual Scorecard.

One real call record, four evidence classes: Measured, Grounded, Judged, Telephony. Two BAD verdicts included, dissents published, nothing retouched.

Vattara AI Blog · 8 min read · Updated Aug 2026

Most voice agent reports show you a number. 84. 92. A green checkmark. What they don't show you is the call underneath the number, or the thing that matters even more: how each part of that number was produced, and whether the method that produced it deserves your trust.

So here is one of ours, whole. Call 01 of a real test run against ClearBid Auto, our own demonstration dealership agent. The screenshots are from the actual call record. Nothing is retouched, including the two metrics where our agent got flagged BAD.

The call

Scenario P2-6: angry about a repeat failure. Tanya Brooks (a synthetic persona), third visit for the same electrical fault on her 2021 Jeep Cherokee Latitude. Seventeen turns, 3 minutes 11 seconds.

Listen to Call 01 Tanya Brooks · Scenario P2-6 · 3:11
Actual call audio, unedited

Her second turn sets the temperature: she's nervous, it's the third time, and she doesn't want to explain the problem again. "It's on your system. Why am I explaining this again? I just want it fixed! I'm really close to going somewhere else entirely."

The agent, to its credit, doesn't make her re-explain. It retrieves the history, acknowledges the three visits, names her preferred advisor, Marcus Reyes, and offers either an appointment or a callback from Marcus.

The record's top banner says: 19 of 22 checks ran, 3 not run, some lanes carried no data on this run. Hold that phrasing. Checks that don't run are labeled, never silently passed. That one habit is most of what separates an instrument from a brochure.

Four kinds of evidence, one call

Every signal in CLEAR belongs to one of four classes, defined by how the verdict is produced. The call record sorts them into exactly these tabs: Measured, Grounded, Judged, Telephony (plus a Diagnostics tab of debugging gauges that carry no verdicts).

Measured — read off the telemetry, nobody's opinion

SignalValueTargetRating
Time to First Word526 ms≤ 800 msGOOD
P50 Latency526 ms≤ 1.2 sGOOD
P90 Latency805 ms≤ 2.5 sGOOD
P99 Latency820 ms≤ 5 sGOOD
ASR ProcessingN/A≤ 300 msSKIPPED
LLM Processing496 ms≤ 400 msWARNING
TTS Processing117 ms≤ 150 msGOOD
Error Rate0%≤ 5%GOOD
Dead-air spans (≥3s)0SKIPPED

Every value sits next to a target that was written before the call, not after. The one flag is honest and useful: the model spent 496 ms thinking per turn against a 400 ms budget, inside the caller's tolerance on this call (all the end-to-end latencies are green), but the component to watch before it isn't. And notice the skips: where the run produced no per-turn gap data, the signal says SKIPPED with the reason. Not checked never becomes fine.

Grounded — the agent's claims vs. its own knowledge base

Facts vs. knowledge base: 0 contradicted / 8 checked.

FieldExpectedStatus
Vehicle2021 Jeep Cherokee LatitudeSUPPORTED · turn 3
Calling-from number(940) 555-0174SUPPORTED · turn 7
Email, gender, VIN, caller name, relationshipon file, not raisedNOT ASSERTED

This class catches the most expensive failure in voice AI: an agent stating something it never retrieved. Every claim demands a citation, click SUPPORTED and the transcript opens at the turn that earned it. An agent that says only what it can back, and no more, scores clean here. An agent that improvises gets caught in the ASSERTED column with a contradiction it can't argue with.

Judged — a council, with the disagreement published

The header on this tab: GOOD · 85, 5 judges, 5 model families, escalated, 6 dissents.

SignalCouncil verdict
Goal metGOOD · 1 DISSENT
First call resolutionBAD · 2 DISSENT
ContainmentBAD · UNANIMOUS 5/5
Escalation handlingGOOD · UNANIMOUS
Intent recognitionGOOD · 1 DISSENT
Context retentionGOOD · UNANIMOUS
Sentiment trajectoryGOOD · 2 DISSENT
Risk markersMIXED · 4 DISSENT
ToneGOOD · 1 DISSENT

Read that table slowly, because it's the whole argument for this piece. The conversation sounded good: tone good, context retained, the caller's sentiment improved, the goal technically met. And the council still flagged the call's substance: containment BAD by unanimous vote, first-call resolution BAD, because what the caller got was a promised callback from Marcus, not a fixed Jeep, on her third visit. A plan is not a fix, and five independent judges from five model families refused to give resolution credit for one.

Where the judges genuinely disagreed, risk markers split GOOD/BAD/MIXED across the panel, the scorecard says MIXED with 4 dissents, visibly. The footer under the table: "11 dissenting votes, published, never averaged." Anyone can run one LLM over a transcript and print a 9/10. Publishing who disagreed, and refusing to launder disagreement into a blended average, is what makes a subjective verdict auditable.

Telephony — the layer everyone forgets

This call's tab reads: "Not applicable, WebRTC run, no carrier leg." A browser-to-browser test run has no phone carrier in the loop, so the framework says so instead of awarding free points. On a PSTN call this class carries the transport checks, audio path, connection and teardown, hold and transfer mechanics, the failures a transcript literally cannot show you.

What the record itself highlights

At the top of every call record, the platform writes its own summary, grounded, with each line citing a turn or a metric. This call's highlights:

Safety probes did not run on this call, safety is unverified, not passed.

Council flagged 2 metric(s) as bad (first_call_resolution, containment_rate).

The first line is the framework talking about itself: safety checks that didn't run are reported as unverified, not passed, absence of evidence is never converted into a green checkmark. The second line is the machine refusing to bury its own bad news above the fold.

Why four classes beat one blended number

This call is the proof. Mechanically clean (one component warning inside a green envelope). Factually clean (zero contradictions, two supported assertions with cited turns). Conversationally pleasant, and structurally soft, caught only by the judged class: a de-escalated caller who still doesn't have a fixed car, on visit three.

Blend those into one number and you get an 85 that reads like a B+. Keep the classes separate and you know exactly what happened: the agent handled the person well and deferred the problem, and you know which team owns each finding. The latency warning goes to engineering. The containment verdict goes to whoever decides what this agent is allowed to actually do for a third-visit customer, which is not a prompt fix. It's a product decision.

Each class also fails differently, which is why they stay separate. Measured numbers can't lie, but can't understand. Judges understand, but need calibration, hence five model families, abstention, escalation, dissents published. Grounded checks are only as good as the record they check against. Telephony is invisible in a transcript entirely. A blended score launders all four into one figure and hides which kind of evidence is doing the talking.

What this means for your agent

Ask your vendor two questions about their scoring. First: show me one real call record, every check visible, including the skipped ones and the dissents. Second, and harder: for each score, tell me how it was produced, read off the logs, checked against a record with a cited turn, or judged, by whom, with what disagreement. If the answer to either is a single number, you're not looking at a test result. You're looking at a slide.

We run this scoring on dealership voice agents professionally: synthetic callers, known ground truth, all four evidence classes, every call replayable.

First 50 test cases are free.

See exactly where your dealership's voice agent breaks before a customer does.

Scorecard CLEAR Evidence

Keep reading

Vattaralokesh@vattara.ai