GPT-5.6 Sol Just Cracked 6 Erdős Problems and a 30-Year Graph Theory Conjecture
OpenAI’s newest flagship model is tearing through decades-old open math problems, but the real breakthrough is the adversarial multi-agent workflow driving it.
The math community is reeling this week as OpenAI's newly released GPT-5.6 family systematically dismantles decades-old open problems. But while the raw reasoning power of GPT-5.6 Sol is staggering, the real story is how researchers are using adversarial multi-agent workflows to turn the model into an autonomous mathematician.
In the span of just a few days, AI has falsified a 30-year-old graph theory conjecture, solved the notoriously difficult Cycle Double Cover Conjecture, and cracked six open Erdős problems. We have officially entered the era where AI is no longer just assisting mathematicians—it is out-researching them.
The Fall of a 30-Year Graph Theory Conjecture
The first major domino to fall was the Dinitz-Garg-Goemans conjecture. Proposed three decades ago, the conjecture posits that if a feasible fractional flow satisfying capacity constraints exists, there must exist an indivisible flow whose capacity is exceeded by at most the value of the maximum demand, without increasing the overall cost. It has been a stubborn "cold case" in graph theory, listed as an open problem in academic literature as recently as January 2026.
Dmitry Rybin, an Olympiad Gold Medalist and AI startup founder, used the new GPT-5.6 Pro model to falsify the conjecture overnight. The model constructed a brilliant, highly structured counterexample involving a triangle graph with three terminals (with demands of 15, 10, and 15).
The AI proved that any valid integer solution requires a total cost of at least 60, because the three "cheap" paths are pairwise conflicting. However, the fractional flow can utilize all three cheap paths simultaneously at ratios of 1/3, 2/5, and 1/3, resulting in a total cost of only 58.
58 < 60. With that simple inequality, a 30-year-old conjecture was dead.
The most shocking part of this breakthrough? Rybin guided the model using only three prompts:
- "Construct a counterexample. You need to make a breakthrough and find a structured counterexample."
- "Keep searching. Develop a clear strategy derived from a deep understanding of the problem's structure."
- "Partial results are enough. Let's directly present a complete, unconditional counterexample."
6 Erdős Problems in 5 Days
While Rybin was dismantling graph theory conjectures, Shouqiao Wang (Qiaoqiao), a PhD candidate at Columbia University, set his sights on the legendary Erdős problems. Named after the prolific Hungarian mathematician Paul Erdős, these are a famous collection of open questions, many of which carry cash bounties for their solutions.
Using GPT-5.6 Sol paired with OpenAI’s Codex workflow, Wang attempted 13 open problems. In just five days, he successfully generated proofs for six of them.
"Solved" is a heavy word in mathematics, and the community will need months to formally verify these proofs. But going 6-for-13 on historically stubborn problems in less than a week is an unprecedented flex of AI reasoning.
The Secret Sauce: Adversarial Prompting and Workflows
The interesting part of Wang's achievement isn’t just the underlying model—it’s the workflow. This wasn't a case of zero-shot magic where a user simply asks a question and gets a Nobel-worthy answer. Wang engineered a brutal, adversarial environment for the model to operate in.
Here is how he structured the winning workflow:
- Strict Boundaries: The prompt restated the problem precisely and explicitly listed what wouldn’t count as a valid proof. Near-misses, weaker results, or redirects to other open problems were strictly forbidden.
- Self-Sabotage: Wang instructed GPT-5.6 Sol to actively hunt for counterexamples to its own lemmas. If a logical path only led to another unsolved problem, the model was told to kill the thread immediately rather than wasting compute.
- Adversarial Agents: Once a proof survived the first pass, a separate layer of adversarial AI agents was deployed to aggressively try and break it. Only proofs that survived this automated "murder board" were considered complete.
Problem selection was also highly strategic. Wang deliberately avoided famous conjectures where the "blast radius" of being wrong is enormous. Instead, he targeted problems that mathematicians actively argue about—contested enough to be interesting, but scoped enough to be tractable for a Large Language Model.
Sol Ultra and the Cycle Double Cover Conjecture
If that wasn't enough, OpenAI's heaviest hitter—GPT-5.6 Sol Ultra—tackled the Cycle Double Cover Conjecture on July 10, just one day after general availability.
The GPT-5.6 family includes three tiers: Terra, Luna, and the flagship Sol. Sol Ultra is the parallel multi-agent reasoning mode, designed for the most complex cognitive tasks. The Cycle Double Cover Conjecture asks whether every graph without a "bridge" edge can have its edges covered by a collection of cycles where each edge appears exactly twice.
Using 64 parallel subagents, Sol Ultra generated a complete proof in under one hour.
OpenAI has made the prompt and the resulting proof publicly accessible, inviting peer scrutiny rather than hiding behind closed doors. Sol Ultra, which scored a massive 91.9% on the Terminal-Bench 2.1 advanced reasoning benchmark, is currently priced at $5 per million input tokens and $30 per million output tokens. While expensive compared to standard models, it is remarkably cheap for a system capable of generating novel mathematical proofs.
The "Last Fields Medal" Era?
Rumors are already circulating in academic circles that the 2026 Fields Medals might be the "last awarded exclusively to humans." While that might be hyperbole, the events of the past week prove that we have crossed a critical threshold in artificial intelligence.
AI is no longer just a calculator, a coding assistant, or a glorified search engine. When paired with adversarial agentic workflows, frontier models like GPT-5.6 Sol are capable of genuine mathematical discovery. The bottleneck is no longer the AI's reasoning capability—it's the human community's ability to peer-review its output.
For AI engineers and developers, the takeaway is clear: the future belongs to those who can build the best scaffolding. The models are smart enough to solve the world's hardest problems, provided you know how to build an arena that forces them to prove it.
Sources
Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.