On September 5, 2026, Anthropic announced that dozens of Claude agents, working largely autonomously over 11 days, completed the first fully formalized, machine-checked Lean 4 proof of Fermat’s Last Theorem. The project generated approximately 13 million lines of Lean code, proved around 30,300 computer-verifiable theorems, and consumed roughly six billion output tokens.
To be precise about what happened: Claude did not solve Fermat’s Last Theorem. Andrew Wiles solved the problem in 1995. What Claude did was formalize the proof — translate Wiles’s mathematical argument into a language a computer can verify step by step, with no gaps and no room for human error. Mathematicians considered this a project that could take many years. Multiple Claude agents completed it in under two weeks.
What Formalization Actually Means
When a mathematician writes a proof, they rely on shared conventions, implicit steps, and the judgment of expert readers. A formal proof eliminates all of that. Every logical inference must be spelled out and verified by a proof assistant — in this case, Lean 4. This means no hidden assumptions, no glossed-over steps, and no errors that slip past human reviewers.
The Fermat formalization is significant because Wiles’s original proof is notoriously complex, drawing on deep areas of number theory, algebraic geometry, and modular forms. Previous attempts at formalization — by human teams — had stalled due to the sheer volume of work required to fill in every detail.
Claude didn’t do this alone. The project used Prove2ME, an open-source tool from Columbia University designed for large-scale mathematical formalization. Prove2ME breaks a large proof into interconnected smaller theorems and tracks the dependency structure between them, allowing multiple AI agents to work in parallel on different pieces without stepping on each other.
Notably, the first attempt failed. Prove2ME was added mid-run after the initial approach hit a wall. That pivot — recognizing failure and changing strategy — enabled completion.
Why This Matters Beyond the Math
This is not just a story about a famous theorem. It is a data point about what AI agents can do when given hard, long-horizon intellectual work.
A few things stand out:
Scale of autonomous operation. Eleven days, largely without human intervention. The agents coordinated, hit dead ends, and recovered. This is qualitatively different from a model completing a task in a single session.
Collaborative agent architecture. Dozens of Claude agents working in parallel, each on a piece of a much larger whole. This mirrors how enterprise teams are starting to use agentic AI — not one agent doing everything, but many agents dividing complex work.
Verification as output. The result isn’t a text summary or a proposed answer — it’s a machine-verified artifact. Every claim has been checked by a proof assistant. That kind of guaranteed correctness is increasingly valuable as AI is deployed in high-stakes settings.
For businesses thinking about AI adoption, the lesson is not “AI can prove theorems.” It is that AI agents are now capable of sustained, high-complexity, multi-step intellectual work over extended periods — and that with the right architecture, the output can be independently verified.
What This Means for Business
The Fermat formalization is a research milestone, but the capabilities it demonstrates have direct business implications.
For firms in finance, engineering, and science: Formal verification is already used in safety-critical software and chip design. The ability to deploy AI agents that can assist with or automate formal verification work represents a meaningful capability jump. Work that previously required rare, expensive specialists can now be approached differently.
For any organization building AI-assisted workflows: This story validates the multi-agent approach. Complex problems can be decomposed, parallelized, and recombined. The architecture Anthropic used — many specialized agents coordinated through a dependency-aware tool — is the same pattern that enterprise AI deployments are moving toward.
For leaders evaluating AI capability: If Claude agents can coordinate autonomously for 11 days on a problem experts thought would take years, the ceiling for what AI can contribute to complex, long-running business problems is higher than most organizations are planning for.
The practical question for most businesses is not whether to use AI for mathematical theorem proving. It is whether they are thinking ambitiously enough about where sustained, collaborative AI work could create value in their operations.
Enterprise DNA’s Omni advisory service helps business leaders think through exactly this kind of question — where the real leverage points are, what architectures make sense for their context, and how to move from one-off AI experiments to production workflows. If you are trying to figure out what the next level looks like for your organization, book a discovery call with Sam McKay.
Source
Anthropic Research
Free Resource
Going deeper with Claude?
Get the free 32-page implementation guide for ANZ teams.
Your guide is ready
Check your downloads folder. If it did not open automatically, use the button below.
Download the GuideWant this working inside your business?
See what's possible