I run a local model and I pay for cloud models, and the most common question I get is "which one should I use?" The honest answer is both, on the same task, at different stages. After a year of building tools that use Ollama and Claude together, here is the decision framework I actually apply, updated for the mid-2026 landscape where the top cloud tier now costs $50 per million output tokens.
The cost gap got wider, which makes the question sharper
In 2026 the cloud frontier spread out into a clear ladder:
Haiku 4.5: $1 / $5 per million tokens
Opus 4.8: $5 / $25






