Back to Articles
The architecture: four ideas carry it Kimi Delta Attention · the workhorse Gated MLA · the precise quarter Attention Residuals · the new one Stable LatentMoE · the sparsity bet The sparsity race, on one chart Benchmarks What it will look like on hfviewer Related links This Hugging Face edition adapts our original interactive Kimi K3 preview at hfviewer.
Moonshot AI announced Kimi K3 on July 16: a 2.8-trillion-parameter open model with a 1-million-token context window, native vision, and two architectural bets that go beyond scaling up the usual recipe. The weights land on Hugging Face by July 27. This page is our deep-dive, prepared in advance: the benchmarks and the architecture story now, the traced graph the moment the checkpoint drops.
The architecture: four ideas carry it
K3 is not a scaled-up K2. Moonshot credits the jump to roughly 2.5× better scaling efficiency than Kimi K2, and the launch post names the mechanisms plainly. Each one lands on a trend line we have been tracking across the 2,400+ models traced on hfviewer.











