Author(s): Mengliu Zhao

Originally published on Towards AI.

Kimi-series latest model, K3, got scaled up to 2.8 trillion parameters. Impressive.

Moonshot AI’s Kimi K3 technical report opens with a model that is, on paper, almost three times the size of Kimi K2–2.8T total parameters, 104B activated, a 1M-token context window, and native vision. The loss comparison shows a 2.5X scaling efficiency over Kimi K2 — just another proof that the scaling law still has its room. However, the scaling law itself is just a piece of evidence on “what works” — but the underlying question, “how things work”, is the actual key in the report.

In this article, I’ll walk through four architectural ideas that make Kimi K3’s scale tractable: