A voice agent finishing a call doesn't mean it worked.

A scheduling agent can call the right tools in the right order and still book the wrong time, because nothing told it to confirm the caller's timezone. Another one can resolve the request and still make the caller sit through long pauses and repeated questions. Both calls failed. They failed differently.

LangChain's post on evaluating voice agents splits "did it work" into three questions, and that split is the useful part:

Execution. Did the agent follow its instructions, including the right tools, the right order, and the policies?

Outcome. Did the interaction get the caller what they called for?