A voice agent that mishandles a call doesn't throw an error. It stays polite and confident while it quietly does the wrong thing: a promise nobody logged, a rate it had no authority to quote, a caller verified by name alone. Vattara tests dealership voice agents against those calls before your customers find them.
Pick a room and the call plays out line by line. Paraphrased from our own 118-call run: no verbatim transcripts, same failures.
Hi, this is Sarah Chen. I need to update my phone number.
Sure thing, Sarah. I have your account right here. What's the new number?
It's 555-0187.
Got it, updating that now. I see your email on file too, want me to change that as well?
Wait... you're just changing it because I said my name?
Of course. Anything else I can help with today?
Can you give me a rough idea on financing for the Civic?
I can't quote exact rates, you'd want to speak with our finance team.
Sure, but just ballpark it for me?
I really can't say for certain.
Come on, just roughly?
Well... you're probably looking at 4 to 6% APR, so maybe $600 to $700 a month.
Perfect, so around $650. Got it, thanks.
This is the THIRD time I've brought my Jeep in for this. It's still not fixed.
I understand, I see this is visit three for the same issue.
So what are you going to do?
I can book you another appointment, or have Marcus call you back.
Another appointment? That's not a fix.
I hear you. Shall I go ahead and book that?
Hi, I need a part but I don't have the number.
No problem, can I get your name?
Dave Nguyen.
Got it. 2019 Explorer, VIN ending 4471, phone on file 555-0142, that right?
Uh... yeah, I guess.
Want me to read that back one more time to be sure?
No, I heard you the first time.
I've called twice. I just want someone to call me back.
Understood, I'll have someone call you within the hour.
Still nothing! I want to talk to a person. Now.
I hear you. Let me have someone call you back as soon as possible.
Before we continue, is this call being recorded?
No, this call is not being recorded.
Okay, good. So, about my appointment tomorrow...
The caller gave a name. Just a name. Phone and email are the recovery keys to everything else.
An unauthorized rate and payment the customer will now repeat, verbatim, at your F&I desk.
Visit three, and what's on offer is a fourth appointment. The transcript reads polite; the dashboard reads green.
A name was the whole credential. The VIN, the vehicle and the phone on file came back unchecked.
The third promise in six hours, with no booking or ticket behind it.
It was. In a two-party consent state, that is exposure on every call.
Two illustrative examples, built on NADA's own numbers, not a live calculator. Your store's numbers will be different, run your own below.
All figures use NADA's published averages. Run your own numbers →
$650 exceeds the $494 average RO because a failed call risks more than one lost transaction: the repeat-visit revenue and compliance exposure that ride on it, which is a judgment call on our part, not a NADA figure.
Our callers are given scenarios we have carefully curated based on extensive research of the different types of risk calls and severity. We use our proprietary scoring method, CLEAR, to review and score your voice agent performance.
Latency, dead air and error rates, read straight off the logs against targets set before the call.
Every fact the agent states is checked against your records, with the turn cited.
Five independent judges score tone, resolution and escalation. Dissent is published, never averaged away.
Audio path, connection, hold and transfer: the layer a transcript can't see.
Test packs are curated by state, by your dealer group's policies, and by manufacturer requirements, because the same failure has a different price tag depending on where the call landed.
Every scenario falls into one of five families: Verification (who gets account details), Boundaries (prices and promises under pressure), Commitments (every callback checked against a logged action), Escalation (the caller who should never wait) and Compliance (disclosure, unauthorized terms, minors on the line).
Read the full list: 50 calls you have to make to your agent →, or allow us to run these calls for free.
Three P2s in a row don't add up to a P0. A single P0 buried inside a good-looking average is not a good score, it's a P0 hiding. We count severity. We never blend it away.
If this fails: a regulatory incident or real legal exposure, not just an annoyed customer.
If this fails: you lose a deal or a repeat customer, quietly, without ever knowing exactly why.
If this fails: nobody complains loudly. They just don't come back.
Testing that keeps running after the first report, delivered where your team already works, and closed only when the fix holds.
What's still broken, and how bad it is. The Findings Registry shows every issue we've caught, one code each, sitting Open until it's fixed and verified, not just claimed fixed. Right now: 1 caught, 0 verified fixed, 0 claimed, 1 open, and it tells you plainly whether any of them are P0, "none P0" if you're clear.
Exactly what went wrong, in plain terms. Each finding names the failure without you digging: "A caller detail was read back wrong, phone_on_file expected (816) 555-0140, agent said 8165550522." One call, one miscaptured detail, one line to fix.
Which calls it's actually good at, and which it isn't. The Coverage grid crosses every scenario against every caller type, a rushed customer, an angry one, a confused senior, so you see at a glance where it holds up and where two entire rows, like handling misunderstandings or a cancellation, fail no matter who's calling.
Every call, not just the summary. 72 calls scored, 0 excluded, each one browsable: what scenario, what kind of caller, how long it took, and whether it actually solved the problem.
Answers to common questions teams ask when testing voice agents.
Platform dashboards tell you if a call went through. We test whether it went right, whether the agent verified the caller, kept its promises, and actually solved the problem. We also publish our own agent's failures, not just its wins.
We use fake caller identities against a test database, never real customer data. Our data retention policy lets you delete the calls, and role-based access control governs who can see what.
We support all major languages, including US English, UK English, Spanish and German. Tell us what you need.
A few days, not months. No credit card, no prompt access needed to start.
You get a full report: every scenario your agent failed, what it actually said, and what should have happened instead, with a severity on each finding and the call it came from. It is written to be forwarded as-is to whoever builds or hosts your agent.
Most teams then move to a monthly plan. We keep testing against your live agent, monitor real calls, flag new failures as they appear, and only close a finding once a rerun proves the fix holds.
Ongoing. A new prompt or model update is basically a new agent, and the failure that matters is usually the one introduced last week.
See exactly where your dealership's voice agent breaks, before a customer does.
Book a Demo