The agent worked. The architecture didn't.

This failure mode won't show up in your eval suite.

An AI architect builds a personal automation pipeline: an agent scouts trending content, pipes ideas into a chat thread, he reacts, things get logged. The agent did its job. It found the trends, wrote them down, and responded when he asked it to.

He killed it anyway, and not over hallucination or token cost. He killed it because everything lived in "a flat scroll of messages. No structure, no views, no way to see what is in scripting versus what is scheduled." He was spending his entire one-hour creation window scrolling backwards through a Telegram thread, looking for something he'd written the day before.

The model wasn't the problem. The substrate was, and almost everyone shipping agents today has the same problem, because the default agent UI is a chat window and the default system of record is the transcript.