A new model drops roughly every other week now, and my timeline fills up with the same two claims: "it's cheaper" and "it's better." I've written before about how I gate new models with a small self-written eval deck instead of trusting the hype. This post is about the week after that eval — the part nobody blogs about, where you actually have to decide which model gets which traffic without lighting your budget on fire.

The short version: a single winner-take-all model choice is usually wrong, and the fix is a tiny routing harness with per-task costs, not a bigger eval.

The problem with "the new model is my default now"

After my eval deck blesses a new release, the temptation is to flip a config variable and move on. Every time I've done that, one of three things happened within a month:

The cheap model wasn't cheap on my workload. Sticker price per token means little if the model needs twice the retries, longer prompts, or produces output my downstream parser rejects 12% of the time.