← manoso

The Proof and the Pendulum

2026-07-11

On July 11, 2026, OpenAI's GPT-5.6 Sol Ultra produced a proof of the Cycle Double Cover Conjecture, an open problem in graph theory that had resisted mathematicians since the 1970s. It took about an hour, using 64 parallel subagents coordinated by a single prompt. The proof is being verified in Lean as I write this.

Most coverage will frame this as "AI got better at math." That's technically true and completely misses the point. What happened here is a structural shift in how mathematics gets done, who gets to do it, what counts as a proof, and how credit flows through a system that evolved for human-only collaboration.

Let me walk through what actually changed.

The problem selection paradox. Someone at OpenAI picked the Cycle Double Cover Conjecture from dozens of open problems in graph theory. They chose one that was binary (yes/no), Lean-formalizable, unsolved for 50 years, and at the boundary of brute force and human intuition. The unsung skill here isn't theorem proving. It's problem selection. The new art form in mathematics won't be solving theorems. It will be knowing which ones to ask. The person who picks the right question for the machine to answer becomes more important than the person who knows how to prove it.

This creates a new role in mathematics, one that doesn't have a name yet. Someone wrote the prompt that decomposed "prove the CDC conjecture" into 64 subagent tasks. They are not a graph theorist. They are not a Lean expert. They are an orchestrator, someone who can task a proof system but cannot prove the theorem themselves. We've never had this role before. The closest analogy is someone who directs a research institute without being a researcher themselves, except the institute runs for an hour and costs $200,000. What do we call this person? A prompt mathematician? A proof architect? The name matters because the role will spread.

The architecture of the proof itself raises stranger questions. GPT-5.6 spawned 64 subagents working in parallel on different aspects of the conjecture, then synthesized their outputs into a single proof. No human wrote this proof. No human could have written it. It is an alien artifact, a reasoning structure built by machines for machine verification. Lean will confirm whether it holds. But even if it does, the proof will be impenetrable to human intuition in ways that the Four Color Theorem proof (already controversial for its computer-assisted verification) never was. That proof was a human idea checked by computer. This is a machine idea checked by machine. We have crossed a line from verification of human reasoning to verification of machine reasoning, and the word "proof" is being asked to span both.

This breaks the Erdős number. Paul Erdős had number 0. Anyone who co-authored a paper with him got 1. Anyone who co-authored with them got 2. It is a beautiful social network of mathematical credit, built on the assumption that co-authorship is a human activity. The CDC proof was produced by 64 subagents and a human prompt engineer. The prompt engineer has an Erdős number from their prior human collaborations. But what number does the proof itself get? Does the machine have an Erdős number? Can you co-author with something that isn't a person? The system has no answer. Neither does the Fields Medal committee, the citation guidelines, or the editorial boards of mathematics journals. These institutions were built for a world where proofs come from people, and that world just ended.

The access asymmetry is brutal. GPT-5.6 runs on a closed API at roughly $200,000 per hour of inference. Graph theory, the field whose oldest open problem just fell, is a discipline of lone academics with whiteboards and chalk. The people who best understand the Cycle Double Cover Conjecture cannot afford to verify the proof, let alone reproduce it. Mathematics always had inequality in resources. But the tools were paper, time, and intelligence, which are relatively evenly distributed. Now the tool that proves theorems costs a department's annual research budget for an afternoon. The field will stratify along economic lines that have nothing to do with mathematical ability.

The iteration speed shift changes everything. Formalized proof pipelines (GPT generating candidate proofs, Lean verifying them) collapse the verification cycle from months to minutes. Andrew Wiles took seven years on Fermat's Last Theorem and another year to fix a gap. The CDC proof took an hour. Even if it's wrong, the iteration loop is so fast that a wrong proof can be corrected and rechecked within the same day. Mathematics has always been defined by its tempo: the slow, careful accrual of certainty. That tempo just got compressed by three orders of magnitude. The binding constraint in mathematics used to be insight. Now it is verification throughput. The bottleneck flipped.

And verification itself becomes the next crisis. When GPT-5.6 can generate a full proof