AI agents that call real tools — deleting records, sending payments,
editing files — don't usually fail by misunderstanding instructions in
bulk. They fail by issuing one bad call in an otherwise-correct session.
An eval score of 98% is no comfort if you're the run in the 2%, and no eval
can tell you, in the moment, whether the call on the stack right now is






