On July 31, DeepSeek dropped a model that sent the usual corners of the internet into a familiar frenzy. DeepSeek V4 Flash 0731 scored 89.0% on ARC-AGI-1 and 61.4% on ARC-AGI-2, according to the ARC Prize Foundation’s verified results page. At $0.14 per million input tokens, the price-to-performance ratio looked extraordinary. Hacker News lit up. The tweets wrote themselves. Another brick in the wall of AI commoditization, another proof point that the Chinese labs are closing the gap, another reason the proprietary frontier doesn’t matter.

There’s just one problem. The gap isn’t closing. It’s stretching.

The Scoreboard Tells a Different Story

One week before DeepSeek’s release, on July 24, Anthropic published Claude Opus 5. Its ARC-AGI-1 score: 97.5%. Its ARC-AGI-2 score: 90.4%. That is not a rounding error. On the harder benchmark — ARC-AGI-2, designed explicitly to resist the memorization and pattern-matching that inflate scores on ARC-AGI-1 — Claude Opus 5 leads DeepSeek V4 Flash by 29 percentage points. OpenAI’s GPT-5.6 Luna, released July 30, scored 90.7% on ARC-AGI-1 and 59.6% on ARC-AGI-2. DeepSeek edged past OpenAI on the harder test by less than two points while trailing Anthropic by nearly thirty.

The celebration around DeepSeek’s numbers is not wrong on its own terms. An 89% on ARC-AGI-1 is genuinely impressive. A 10-point jump on the Artificial Analysis Intelligence Index over the previous V4 Flash, released just three months earlier, is real progress. But the framing — the implicit claim that this represents convergence — collapses the moment you look at the leaderboard as a whole. The frontier is not standing still waiting to be caught. It is accelerating.

Commoditization Is the Story We Tell Ourselves

There is a reason the commoditization narrative has such staying power. It flatters several audiences at once. For the open-source community, it validates the thesis that weights want to be free. For enterprise buyers, it promises that today’s eye-watering inference bills are temporary. For policymakers, it suggests that AI dominance is not winner-take-all and that export controls might be working just fine, thank you. Everyone gets to feel like a clear-eyed realist while the numbers tell a messier story.

A research engineer at a midsize AI lab put it plainly in a Slack DM after the ARC results dropped: “We’re all clapping for the guy who ran a four-hour marathon while the Kenyans are already showered and on the bus.” The metaphor is imprecise — DeepSeek is running closer to 2:30 than 4:00 — but the dynamic it captures is real. The frontier labs are not merely ahead. They are ahead by margins that are growing in absolute terms, even as the追赶者 improve.

Look at the ARC-AGI-2 scores in sequence. Claude Opus 5 at 90.4%. DeepSeek V4 Flash at 61.4%. Gemini 3.6 Flash at 60.4%. GPT-5.6 Luna at 59.6%. Grok 4.5 at 52.6%. The spread between first and fifth place is 38 points. That is not a tightly bunched field. That is a hierarchy with a clear leader and a pack fighting for scraps.

What Happens When the Frontier Accelerates

The practical implication is uncomfortable. If the gap on ARC-AGI-2 is a reasonable proxy for generalized reasoning capability — and the ARC Prize Foundation designed it precisely for that purpose — then the models most people will actually use, the cheap ones, the ones that run on consumer hardware, are not catching up to the frontier. They are falling behind it on an absolute basis, even as their scores improve year over year.

This matters because the economic value of AI does not distribute evenly across the capability curve. The models that can reliably solve novel reasoning problems — the Claude Opus 5 tier — unlock use cases that the DeepSeek tier cannot touch. If that capability gap widens rather than narrows, the market does not commoditize. It stratifies. A small number of labs capture the high-value reasoning work. Everyone else competes on price for summarization, classification, and code completion — useful, but not transformative.

None of this is an argument against DeepSeek. The model is a remarkable engineering achievement, and the price point is genuinely disruptive for a wide range of workloads. But the story we tell ourselves about what that achievement means deserves scrutiny. A rising tide does not lift all boats if some boats are rising twice as fast. The ARC-AGI leaderboard, read honestly, does not show convergence. It shows a frontier pulling away from the field — and a field that seems increasingly content to celebrate its own progress rather than measure the distance to the lead.

Sources