AirLLM's README opens with a line that sounds like it can't be true:
AirLLM dramatically reduces inference memory usage, letting 70B large language models run on a single 4GB GPU card — without quantization, distillation, or pruning.
So let's actually check it. Is the claim real? How do you set it up? And — the question nobody asks loudly enough — should you?
TL;DR: The claim is technically true and the engineering is legitimate.
The trick, in one paragraph






