How to Run an 80B Qwen Model in 4.3GB of RAM: The Edge AI Revolution Explained

It started with a single Hacker News post — a screenshot of system_profiler showing 4.3GB of memory used by Qwen 80B, running at an uncomfortable but usable 4 tokens per second. Within hours, someone posted a follow-up: a 35B model running on an iPhone 18 Pro, not in the cloud, not even in the high-end Pro Max, but the base model. The thread exploded. Skeptics called it clickbait. Then the benchmarks arrived.

Welcome to 2026, the year edge inference stopped being a trade-off between size and practicality.

The 4.3GB breakthrough: It's not just quantization

When the first reports appeared, the immediate assumption was that someone had used a 2-bit quantization to squeeze an 80B model into tiny memory. That's true — but it's only part of the story. Modern quantization has evolved beyond simple weight rounding.