Originally published on tamiz.pro.

For the better part of two years, the narrative has been seductive: build a few function calls around a reasoning model, add a tool-use loop, and you have an autonomous agent that solves hard problems. The demos are real. The benchmarks are impressive. And then you try to ship one to production — and the floor opens up beneath you.

I've spent the last eighteen months building, breaking, and rebuidling AI agent systems across multiple teams and product lines. What follows isn't a critique of the technology. It's a map of the gap between what was promised and what actually works at scale, written for engineers who are currently standing in that gap.

The Evaluation Illusion

The first hard lesson is that every benchmark you see published is measuring something fundamentally different from what your production system does.