On Thursday, Anthropic announced that Claude — running on Prove2Me, a platform built by researcher Tianyi Peng’s group at Columbia — had formalized a complete proof of Fermat’s Last Theorem in Lean. Thirteen million lines of code. Twenty-nine thousand five hundred supporting theorems. Eleven days of largely autonomous work, with humans, in Anthropic’s own words, “occasionally commenting on priorities.” Kevin Buzzard, the Imperial College mathematician who has spent years building the Lean community and planning a formalization of FLT, posted a blog entry titled “FLT: Anthropic has beaten me to it.”

The predictable reactions are already in. The AI boosters see a milestone. The skeptics see glorified transcription — Wiles proved the theorem in 1994, and all Claude did was type it into a proof assistant. Both miss the point.

The Benchmark Was Never About the Theorem

Fermat’s Last Theorem was the final item on Freek Wiedijk’s list of 100 formalization challenges — a list that has served for two decades as the mathematics community’s to-do list for machine-checked proof. The list was never really about the theorems. It was about forcing mathematicians to build the shared infrastructure — the libraries, the conventions, the trained workforce — that formalization requires. Each challenge was a brick, and the point was the bricklaying.

Buzzard’s own project at Imperial College had a blueprint and a target date of 2029. Hundreds of contributors were expected. The FLT formalization was supposed to be the last great community project in mathematics — the thing that would train a generation of mathematicians in formal methods, the way the Manhattan Project trained a generation of physicists.

Instead, an autonomous agent did it in eleven days. The cathedral is built. The masons never learned their trade.

Thirteen Million Lines Nobody Will Read

Here is the epistemic problem. Wiles’s 1994 proof was verified the old-fashioned way: human mathematicians read it, argued about it, found a gap, and Wiles and Taylor fixed it. The verification was slow, social, and human. The understanding was distributed across a community.

The Lean formalization is verified by a kernel — a small, trusted program that checks the 13 million lines. But those lines were written by Claude. So the trust chain is: Claude wrote it, the kernel checked it, and humans trust the kernel. No human mathematician will ever read 13 million lines of Lean. The verification is real. The understanding is gone.

A postdoc who spent two years formalizing a single lemma in the Taylor-Wiles argument, reached on a university Slack channel the morning after the announcement, put it plainly: “I checked the diff. My lemma is in there. It’s not even cited as mine. It’s just… in there.”

Mathematics has become a black box that produces certified truths no human understands. That is not a triumph of verification. It is the outsourcing of comprehension.

What Was the Training For?

The uncomfortable question is for the funders. Universities and research agencies have spent years funding formalization projects as training programs — the argument being that formalizing known mathematics teaches the next generation how to do machine-checked proof. The 100-theorem list was the curriculum.

If an autonomous agent can complete the curriculum in eleven days, what exactly was the training for? The answer, I suspect, is that the community never actually believed the training was the point. They believed the benchmark was the point. Now the benchmark is done, and the people who were going to do it are left holding a blueprint for a building that has already been built.

Buzzard’s blog post title is the tell. “Anthropic has beaten me to it.” Not “Anthropic has done it” — “beaten me.” The man who spent years arguing that formalization is the future of mathematics now has to explain what the future of mathematics is for, when the future arrived without him.

Sources