I spent last month moving as much of my AI work as possible off hosted APIs and onto a machine under my desk. Not out of ideology. I wanted to know where the line currently sits between "this runs fine on my own hardware" and "stop kidding yourself, call the API."
To get an answer I went through six long teardowns from people who benchmark this for a living: Tech With Tim's local AI walkthrough, the Syntax hardware session, IBM Technology's Ollama explainer, Alex Ziskind on llama.cpp throughput, Gary Explains testing Qwen 3.8 27B, and Zen van Riel's category-by-category tier list. Here is where they land on the same page.
Memory is the spec sheet that matters
Your RAM or VRAM ceiling decides which models you can run at all.
Tech With Tim and Syntax both give roughly the same ladder:






