"Just self-host an open model, it's free." I've heard this a lot, and it's one of the most expensive sentences in tech. Open weights are free. Running them is not. Here's the honest cost comparison between hosting your own LLM and just calling an API, with real 2026 numbers.

The short answer

For most teams, calling an API is cheaper, and it stays cheaper until you're pushing serious volume, roughly 5 to 10 million tokens a month on a premium model. Past that, self-hosting can win on raw compute, but hidden costs like engineering time often erase the savings anyway. Scale and specific needs decide it, not vibes.

How API pricing works

The API model is simple: you pay per token, and someone else owns the hardware. Prices in 2026 keep falling. Claude Sonnet 5 runs about $2 per million input tokens and $10 per million output, OpenAI's GPT-5.6 family starts near $1 per million input, and cheaper frontier models like Meta's Muse Spark sit around $1.25 input and $4.25 output.