AirLLM's README opens with a line that sounds like it can't be true:

AirLLM dramatically reduces inference memory usage, letting 70B large language models run on a single 4GB GPU card — without quantization, distillation, or pruning.

So let's actually check it. Is the claim real? How do you set it up? And — the question nobody asks loudly enough — should you?

TL;DR: The claim is technically true and the engineering is legitimate.

The trick, in one paragraph