Another open-weight release, another week of screenshots. This time it's MiniMax M3 filling my timeline, and the pattern is identical to every release before it: polished demos everywhere, reproducible evidence almost nowhere.

I've already covered why benchmark passes shouldn't earn merge privileges, and how to build a multi-day harness for comparing free coding models properly. But there's a cheaper question that comes first: is this model even worth feeding into that harness? Most aren't. What follows is the screening protocol I use to answer that in under an hour, before I invest a single weekend hour.

The output isn't a ranking. It's a short memo: proceed, park, or pass — with receipts.

Why demos can't answer the only question that matters

Launch-week content almost always answers "can this model do something impressive?" Of course it can. Every frontier-adjacent release can. The questions that actually determine whether a model belongs in my workflow are narrower: