The release notes say it's faster. The launch thread says it beats everything. Three people I follow have already switched. And yet, every time I've switched on that basis alone, I've quietly switched back two weeks later after the model mangled a migration script or confidently explained a bug that didn't exist.
The gap is simple: public benchmarks measure what benchmarks measure. My daily work is a pile of half-remembered Django internals, shell one-liners I keep re-googling, and docstrings for functions nobody else will ever read. None of that shows up on a leaderboard. So instead of arguing about rankings, I built a small ritual: when a model worth caring about appears, I spend one evening replaying my own recent work through it. This post is that ritual — the corpus format, the runner, and the scorecard I use to decide.
Ask three separate questions, not one
"Is this model good?" is unanswerable. These are not:
Does it survive my prompt shapes? Mine are ugly: truncated stack traces, two files pasted back to back, instructions like "don't change the public API."






