Your agent didn't get dumber. Its context rotted.
A support agent answers perfectly for 30 turns, then starts getting things wrong. Nobody touched the prompt. The chat just got long — and the one fact that mattered (which card a refund goes to) is now stranded in the middle of a 250k-token window, where the model barely attends to it.
This is context rot, and it's now measurable. Chroma's 2025 study showed accuracy dropping well before the documented context limit, with a sharp knee far under the "1M-token" ceiling — and the unsettling part: coherent, well-structured input can degrade attention more than shuffled input. Standard needle-in-a-haystack benchmarks miss it because the real failures live in long, multi-tool agent sessions.
The fix isn't a bigger window. It's a smaller, cleaner one.
ContextForge is an open-source context compiler. It sits between your app and the model and does four things to everything entering the window:






