For most of the AI coding race, proprietary models held a comfortable lead over their open-weight counterparts. That comfort zone has essentially evaporated.

Code Arena’s WebDev leaderboard, updated on August 6, 2026, shows Anthropic’s Claude Opus 5-max sitting at the top with an Elo score of 1686. Right behind it, at 1675, is Moonshot’s Kimi K3-max, an open-weight model. That 11-point gap is barely a rounding error compared to the roughly 150-point chasm that separated the two categories not long ago.

How the leaderboard stacks up

Code Arena isn’t your typical benchmark suite. The platform relies on blind, pairwise human votes, where real users compare two model outputs side by side without knowing which model produced which result. Those preferences get converted into Bradley-Terry/Elo-style scores, the same rating system used in competitive chess.

The WebDev leaderboard specifically tests models on their ability to build and iterate on frontend web applications, testing models as autonomous agents tackling real-world coding tasks rather than static, multiple-choice benchmarks.