Multi-step reinforcement learning for tool-using AI agents collapses mid-training not because the model loses its skills, but because the probabilities of a few structural control tokens spike and scramble the agent's execution scaffolding. The fix, according to a new paper, Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It, is to interleave supervised learning with the reinforcement training, keeping those control tokens in check.
Key facts
What: A new paper pins the sudden collapse of multi-step tool-use training on runaway probabilities in a few control tokens, and shows that mixing in supervised examples stabilizes it.
When: 2026-06-26
Primary source: read the source (arXiv 2606.26027)







