On Thursday, the ARC Prize foundation published verified scores for DeepSeek’s latest model, V4 Flash 0731. The numbers are striking: 89.0% on ARC-AGI-1, 61.4% on the harder ARC-AGI-2. The model, released July 31, now sits near the top of a leaderboard designed to measure fluid intelligence — the kind of abstract pattern recognition that, until very recently, was considered a safely human advantage.

The reaction on Hacker News was predictable: 561 points, 336 comments, a familiar mix of awe and anxiety. Is AGI arriving faster than we thought? Is China pulling ahead? Should we panic?

Those are the wrong questions. The right one is hiding in the fine print of the results page, and it has nothing to do with capability.

The Number That Matters Is $0.02

DeepSeek V4 Flash 0731 achieved its 89% score on ARC-AGI-1 at a cost of $0.02 per task. The ARC-AGI-2 result came in at $0.04 per task. For context, DeepSeek’s own pricing page, updated August 3, lists the model at $0.14 per million uncached input tokens and $0.28 per million output tokens. These are not enterprise-contract numbers negotiated behind an NDA. They are public, list-price numbers anyone with an API key can pay.

Two cents. Four cents. That is what it now costs to run a reasoning task that, until roughly eighteen months ago, no AI system could solve at all.

This is not a story about benchmarks. It is a story about unit economics. When the cost of a reasoning task collapses to the price of a text message, the business model built around selling reasoning — and that is, at bottom, what every frontier AI lab is doing — begins to look fragile in ways the current revenue numbers do not yet capture.

The Leaderboard Everyone Is Ignoring

ARC-AGI-2 is a genuinely hard test. The average human scores around 66%. The grand prize threshold for the ARC Prize 2026 competition is 85%, with $700,000 on the line. GPT-5.5 hit 85% in April. Claude Opus 4.6 manages 69%. DeepSeek’s 61.4% is respectable but not dominant — it trails the Western frontier by a meaningful margin on the metric everyone is trained to care about.

So why does DeepSeek’s result matter more than GPT-5.5’s? Because OpenAI and Anthropic are not publishing per-task costs on a public leaderboard. They are not forcing the conversation toward price. DeepSeek is.

A pricing strategist at a major cloud provider, speaking on condition of anonymity from a conference room at a Las Vegas tech summit this week, put it bluntly: “Every lab has a model that can reason. The differentiator used to be who had one first. Now it’s who can deliver it cheapest without going broke. DeepSeek just set the floor, and it’s lower than anyone’s internal cost model.”

That is the dynamic the benchmark-obsessed coverage misses. Capability is converging. Cost is diverging. And in a market where the product is increasingly interchangeable, the low-cost producer wins.

What Happens When Reasoning Is a Utility

The AI industry has spent three years convincing investors that intelligence is a premium product — that the models which reason best will command pricing power, build moats, and generate margins that justify trillion-dollar valuations. DeepSeek V4 Flash 0731, priced at fractions of a cent per task, suggests a different future: one in which reasoning is a utility, priced near the marginal cost of compute, and the only durable advantage is scale and operational efficiency.

This is not a new pattern. It is the pattern of every technology market that ever matured. What is new is the speed. The gap between “AI cannot solve ARC-AGI-1” and “AI solves it for two cents” was roughly two years. The gap between “reasoning is a premium feature” and “reasoning is a line item on an API bill” may be even shorter.

None of this means DeepSeek has won. The model still cannot touch ARC-AGI-3, the newest and hardest benchmark, where even the best systems score below 1% while humans score 100%. Genuine, open-ended fluid intelligence in novel environments remains a distant target. But the market does not price stocks on ARC-AGI-3. It prices them on revenue growth, margins, and competitive positioning. And on those measures, a model that delivers near-frontier reasoning for pocket change is more disruptive than a model that wins a benchmark by five points.

The Real Contest Has Barely Started

The ARC Prize was designed to measure intelligence. It may end up measuring something else entirely: who can afford to give reasoning away. DeepSeek just named its price. The question now is whether anyone else can match it — and what happens to the ones who cannot.

Sources