You picked a free model because the answers looked good. Good answers are not an endpoint. An endpoint is the model plus the server plus the network. Demos pass. Pipelines stall. The model was rarely the problem.

So why do we keep benchmarking only the model? Because it is easy. You paste a prompt. You read the output. You declare a winner. The server never gets a vote.

This post is a reproducible benchmark. It measures the pair, not the model. Run it before you wire any free endpoint into CI.

The Pair, Not the Model

Most evaluations compare answers. You paste a prompt. You judge the output. You pick a winner. That measures the model. It ignores the server.