The honest numbers on CPU, NPU, and iGPU inference — what runs, how fast, and when to stop pretending.
Last year a friend who runs a small consulting firm in Dubai asked me what GPU he needed to "run AI locally." He had a budget in mind and had already been quoted a four-figure price tag for a workstation. I asked him what he actually wanted to run. A document chatbot over his firm's contracts. He owned a three-year-old office laptop with 16 GB of RAM and no discrete GPU.
I told him not to buy anything yet. Two hours later, I had a 7B-parameter model running on that laptop at about 11 tokens per second, answering questions from his own PDFs. Not fast enough for a chat product serving a thousand users. Fast enough for a private assistant that costs nothing per query and keeps every contract on his machine.
The market wants you to believe local AI requires expensive hardware. The truth is more interesting and much cheaper. Here is what is actually possible — with real numbers, real architecture, and an honest map of where CPU-only inference stops being viable.
Why Everyone Thinks You Need a GPU








