Field notes on testing, monitoring, and evaluating voice agents, from the team building the CLEAR framework.
A voice agent can work flawlessly in a quiet office but fall apart the moment a real customer calls from a car, a restaurant, a factory floor, an airport, or a busy home.
Monitoring tells you that something is wrong. Observability is the evidence that explains why, across audio, transcript, state, tools, backend, and timing.
Most voice agent test cases only prove a prototype works. They rarely prove it will survive a real caller. The seven-part structure that exposes where conversation and backend logic break.
Five failure points on the invisible pipe between CDK, Reynolds, or Tekion and your agent's mouth: stale data, timeouts, silent writes, expired auth, schema drift.
Real recordings from the same 118-call run. An account takeover that took one name, a recording lie, and a boundary that eroded on the third polite ask.
One real call record, four evidence classes: Measured, Grounded, Judged, Telephony. Two BAD verdicts included, dissents published, nothing retouched.
118 test calls against the same agent, bare and hardened. Six failure families, graded by severity, and why prompt hardening fixes manners but not architecture.
The phone-pain field study: 49 of 101 dealership groups have customers publicly complaining about their phones, in written, permanent, public reviews.
Judgment, authority, or witness: if the call needs one of the three, keep it human. Where the automation line belongs, and how to test the handoff at it.
LLM evals measure the quality of an answer. Voice AI evals measure the quality and outcome of an entire spoken interaction, from the caller’s audio to the business result.
Most teams treat voice-agent reliability as a post-launch support problem. It is an engineering practice, and it starts long before launch day.