I kept noticing the same thing while building agents: when the agent screws up —

picks the wrong tool, sends bad arguments, over-refuses something harmless — that

failure is genuinely useful. It's a concrete example of my model doing the wrong

thing, which is exactly what I'd want to fine-tune away.

And then it just... sits in the logs. Buried in LangSmith/Langfuse exports, mixed