A Wharton professor used GPT-5.6 Sol Pro to disprove, in about 90 minutes, a conjecture that had stood open for thirty years around the Benjamini-Hochberg procedure, a pillar of applied statistics cited more than 130,000 times. GPT-5.5 had failed on the same problem after roughly 20 hours. The counterexample, the code and the full chat logs are public.
Key Takeaways
- Edgar Dobriban (Wharton) got GPT-5.6 Sol Pro to build a counterexample to a statistics conjecture open since 1995
- The false discovery rate provably exceeds the 0.1 target, reaching 0.104 in the model the AI constructed
- GPT-5.5 failed after about 20 hours where Sol Pro concluded in roughly 90 minutes
Have an AI Sum Up This Article
ChatGPTA Thirty-Year Conjecture Falls in One Working Session
The Benjamini-Hochberg procedure, published in 1995 by Yoav Benjamini and Yosef Hochberg, controls the false discovery rate when thousands of hypotheses are tested in parallel. Its founding paper carries more than 130,000 citations, and the model that just shook it had its launch covered when OpenAI opened the worldwide rollout of GPT-5.6.
One assumption had never been proven: the procedure was believed to stay reliable on correlated, normally distributed data under two-tailed tests. Thirty years of literature rested on that unproven trust.
Correlated data is not an edge case. It is the ordinary condition of most large-scale testing, which is why the question kept its status of open problem instead of fading as a technical curiosity. Whoever settled it, in either direction, was going to move the field.
Edgar Dobriban, associate professor at Wharton (University of Pennsylvania), handed GPT-5.6 Sol Pro nothing but the formal definition of the procedure. About 90 minutes later, the model had constructed a counterexample: a setting where the false discovery rate provably exceeds the 0.1 target and reaches 0.104. The whole reasoning chain is checkable, since Dobriban released the preprint together with the code and the complete chat logs of the session.
The contrast with the previous generation sets the scale of the jump. GPT-5.5, handed the same problem, ground away for roughly twenty hours without producing a valid counterexample. The ratio between those two durations does not measure speed: it measures access to a class of problems that stayed locked four months ago.
Berkeley statistician Will Fithian framed the question as the most interesting open problem in his area of statistics. Its resolution by a commercial model in one working session pins a date in the short history of AI applied to research.
A Tiny Gap That Cracks Thirty Years of Trust
The gap between 0.1 and 0.104 looks trivial. Dobriban himself tempered the result: the overshoot stays relatively small and the practical implications still need work. The value sits elsewhere. A guarantee assumed universal just fell, and nobody knows yet how far it degrades in less friendly configurations.
The nature of the result matters as much as its size. The model did not produce a simulation hinting at an overshoot: it constructed a setting where the rate exceeds the target in a provable way. A proof gets checked line by line and closes the debate, where a numerical experiment would have opened years of counter-verification.
The model’s method also clarifies what these systems actually do. Sol Pro invented no new mathematics: it combined existing methods until the counterexample held together. That matches the performance regime we measured when we pushed Sol, Terra and Luna through our GPT-5.6 test: a reasoning depth that scales with the compute budget granted.
GPT-5.6 Sol Pro is precisely the extended-compute variant of the lineup: more reasoning tokens per problem, hence longer mathematical chains. The standard version has been broadly available since July 9. The gap between the two costs money, and this result hands it its first spectacular justification.
For researchers and data scientists, the operational implication is direct: pipelines relying on Benjamini-Hochberg with correlated data deserve a second look. The counterexample provides the template of the unfavorable case to hunt in your own data.
Dobriban’s starting move also sketches a reusable working method: hand over the bare formal definition, without hints or literature, and let the model search. That minimal protocol makes the result hard to dispute on contamination grounds.
More articles on Horizon
- Claude Pro Goes Free for K-12 Teachers in the US
- DeepSeek Raises $1.5B and Plans an IPO at $71B
- Siri iOS 27 Public Beta Opens to Every iPhone
OpenAI Gets Its Research Showcase Against Rivals
For OpenAI, the timing is gold. The company has argued for months that its extended-compute models produce original research, and here is a verifiable, published, reproducible mathematical result obtained by an outside customer with no privileged access.
The transparency of the file sharpens the commercial argument. Published chat logs, open code, accessible preprint: every step of the model’s reasoning can be audited by the community. No staged demo survives that level of exposure, and that is exactly what gives the result its reference value.
Distribution still runs under political constraint. Washington screens access customer by customer, a setup we detailed when the White House started vetting GPT-5.6 customers. A model that knocks down conjectures becomes an argument in that negotiation as much as a product.
On the competitive side, the bar just moved. Anthropic and Google will need to show their own research results attributable to their models, on open problems checkable by third parties. Synthetic benchmarks will no longer settle the extended-compute race.
What comes next hinges on reproducibility. If other academic teams extract counterexamples or proofs on their own conjectures, extended compute becomes a standard budget line in research labs. If the Dobriban case stays isolated, it joins the list of brilliant one-off demonstrations.
Follow the story on Horizon.


