Your model says it has a one-million-token context window. Its real working memory is a lot smaller than that.

On long-context benchmarks, models start failing well below the number printed on the box. And here's the part nobody warns you about: past a certain point, adding more context makes your answers worse. Not better. Worse.

Prefer to watch? Full walkthrough with the attention-cost animation:

A context window is not memory

Let's kill one idea first. A context window is not memory.