A new open model release trends, the benchmarks look incredible, and by the time I've read the third hot take I still have no idea whether it's actually good for my work. The recent chatter around MiniMax's H3 release is a perfect example: lots of excitement, lots of leaderboard screenshots, and very little about how it behaves on the boring tasks I actually do every day.

So I stopped reading takes and built a tiny ritual instead: a 30-minute, reproducible smoke test I run against any newly hyped model before I let it anywhere near a real project. This post is that ritual, plus the exact script.

The problem with launch-week benchmarks

Public benchmarks answer "is this model smart?" I need answers to different questions:

Does it follow my prompt style, or does it need babysitting?