AI agents that call real tools — deleting records, sending payments,

editing files — don't usually fail by misunderstanding instructions in

bulk. They fail by issuing one bad call in an otherwise-correct session.

An eval score of 98% is no comfort if you're the run in the 2%, and no eval

can tell you, in the moment, whether the call on the stack right now is