When DeepSeek released V4 Flash on July 31, 2026, it quietly accomplished something that would have seemed impossible six months ago: it scored 89.0% on ARC-AGI-1 and 61.4% on ARC-AGI-2 — the hardest publicly available abstract reasoning benchmarks — for about two cents per task.

To put that in perspective: GPT-5.2 Pro and Claude Opus 4.8, the most expensive frontier models, cost orders of magnitude more per task on the same benchmark and score only marginally higher. DeepSeek V4 Flash is open-source, weighs in at a fraction of the cost, and runs three reasoning variants (Max, High, Low) that let you trade accuracy for speed.

This is the story of how a Chinese AI lab quietly caught up with — and in some ways surpassed — the entire Western frontier on reasoning, and what it means for developers.

What Is ARC-AGI and Why Does It Matter?

ARC-AGI (Abstraction and Reasoning Corpus) is designed by François Chollet to test something different from traditional benchmarks. Instead of measuring knowledge or language fluency, it tests fluid intelligence — the ability to solve novel visual reasoning puzzles you have never seen before.