Self-hosting Kimi K3 became technically possible on 27 July 2026, when Moonshot AI published the weights for a 2.8-trillion-parameter model alongside production inference support. A great many organisations read that news and concluded they could now run frontier-class reasoning on their own hardware and stop paying per token.
That conclusion is usually wrong, but not for the reason people expect. The engineering is achievable. The arithmetic is what defeats most projects, and it defeats them quietly, several months after the budget was approved.
The short answer: At MXFP4 precision, 2.8 trillion parameters occupy roughly 1.4 TB before any key-value cache. A single eight-way H100 node holds 640 GB and therefore cannot serve this model at all. Realistic deployments start at around 1.7 TB of VRAM, meaning current-generation nodes with 288 GB accelerators or sixteen-way configurations of the previous tier, and Moonshot points production users towards clusters of sixty-four or more accelerators.
What Open Weights Actually Give You
Before the hardware, the licence, because it determines whether any of this is worth planning.






