The race to build better large language models isn't slowing down, and this week Moonshot AI introduced another major milestone: Kimi K3.

At first glance, the headline is impressive enough—a 2.8 trillion parameter open Mixture-of-Experts (MoE) model with support for a 1 million token context window. But after reading through the technical paper, what stood out to me wasn't just the size of the model. It was the engineering behind it.

Instead of simply scaling parameters, Moonshot AI focused on solving several practical bottlenecks that affect training efficiency, inference speed, and long-context reasoning.

Let's take a closer look.

Kimi Delta Attention