I run a local model and I pay for cloud models, and the most common question I get is "which one should I use?" The honest answer is both, on the same task, at different stages. After a year of building tools that use Ollama and Claude together, here is the decision framework I actually apply, updated for the mid-2026 landscape where the top cloud tier now costs $50 per million output tokens.

The cost gap got wider, which makes the question sharper

In 2026 the cloud frontier spread out into a clear ladder:

Haiku 4.5: $1 / $5 per million tokens

Opus 4.8: $5 / $25