Voice agent evals

How to measure whether an AI phone agent actually works — before a real customer finds out it doesn't.

A voice agent eval is a repeatable test that measures an agent's behavior under realistic call conditions before it reaches live callers. Text-agent evals check what the model says; voice evals also have to check when it says it, whether it heard correctly, and whether the call ended in the outcome the business needed. Skipping them doesn't make an agent unevaluated — it makes your customers the eval.

The metrics that matter

Simulated conversations do the heavy lifting

You can't regression-test a phone agent by calling it yourself twice. The working pattern is a suite of simulated callers — scripted personas that interrupt, change their minds, mumble, and go off-script — run against the agent over live audio, with each conversation graded against a verifiable end state. Deterministic checks (was the booking created? was the handoff triggered?) anchor the suite; LLM-as-judge scoring covers tone and policy adherence where no deterministic check exists.

Every deployment we run ends the same way: the eval suite is the deliverable. The agent is just the thing it measures.

Where this fits

Evals are the final stage of the FDE roadmap because they're what separates a portfolio demo from deployment proof. To see the checks applied against real businesses, read the deployment walkthroughs — each one ends with the eval gate that change had to pass.

Go forward, faster.

One email a week on breaking into Forward Deployed Engineering — roles, tactics, and lessons from the field.