Been heads-down lately, but I finally had a couple of days to catch up on some recent papers. Sharing one of them here.
Milestone-Guided Policy Learning for Long-Horizon Language Agents
Zhejiang University (ZJU-REAL) · arXiv:2605.06078 · github.com/ZJU-REAL/BEACON
It tackles a very concrete and fairly painful problem: when you train an agent on long tasks with RL, it doesn't just get a bit worse — it falls off a cliff as the horizon grows.
TL;DR







