I kept noticing the same thing while building agents: when the agent screws up —
picks the wrong tool, sends bad arguments, over-refuses something harmless — that
failure is genuinely useful. It's a concrete example of my model doing the wrong
thing, which is exactly what I'd want to fine-tune away.
And then it just... sits in the logs. Buried in LangSmith/Langfuse exports, mixed






