OpenAI unveiled Astra, its next major model, on August 1 by publishing ten results to mathematics and theoretical computer science problems that had stayed open for at least a decade. Sam Altman demonstrated it to regulators in Washington, but the model is not public: it stays in testing and would have to clear a coming federal approval framework before any release.
Key Takeaways
- Astra produced ten results on problems open for a decade, spanning geometry, lattice cryptography and group theory.
- OpenAI formalized the proofs in Lean and takes responsibility for their correctness, with the arguments generated by the model.
- The model stays in testing, with no release date, and would be the first submitted to the new US federal approval framework.
Ten problems open for a decade
The move is unusual: OpenAI is not showing a benchmark, it is showing results. The company published ten solutions to problems that had seen no progress on their main result for at least a decade, and in most cases much longer.
The fields covered are not trivial. They run from high-dimensional geometry to coding theory, through arithmetic circuit complexity, group theory, quantum complexity, lattice cryptography and extremal combinatorics. Areas where progress usually gets counted in years of human work.
The method matters as much as the result. OpenAI notes the mathematical arguments were generated by the system, while the company prepared the manuscripts and formalized the proofs in Lean, a language that turns a proof into a machine-checkable certificate. It takes responsibility for their correctness.
This is not the first signal of its kind. In May, OpenAI had already shared a model-generated disproof of a longstanding conjecture on unit distances. Showing ten results at once shifts the scale, and moves the story from an isolated flourish toward a capability the company frames as repeatable.
A model built for long-horizon work
Astra is not pitched as a faster chatbot, but as a model built for tasks that stretch over hours, even days. It coordinates several agents working together, a direction that extends the turn already taken with the agent app that replaced the old conversational ChatGPT.
The cost OpenAI describes is strikingly modest. Generating all ten solutions reportedly ran about 2,000 dollars in API rates. The research team leans on that point: each problem used little compute, and there is wide room to push test-time compute much further.
That economics reframes the story. If results at this level fit inside a few thousand dollars, the limiting variable is no longer the price of compute, but the ability to steer a model across a long task without it drifting. That is exactly where OpenAI concentrates its product work, as when GPT-5.5 Instant became the default model to absorb everyday volume.
The competitive read is direct. A model that can hold a scientific line of reasoning across several days targets a market no one really occupies yet, assisted research at scale. Rival labs will have to show not a higher exam score, but comparable endurance on real problems.
More articles on Horizon
- DeepSeek V4 Flash Nears GPT-5.6 Luna at 60% Lower Cost
- GPT-5.6 Luna: OpenAI Slashes the Price by 80%
- Gemini Robotics ER 2 Makes Robots Work Together
Not public yet, and a regulatory gate ahead
The caveat that cools the announcement fits in one word: Astra is not available. The model stays in testing, with no release date announced, and OpenAI has not settled whether it ships as GPT-6 or as an intermediate variant.
A new obstacle sits upstream. Astra would be the first model submitted to a new US government framework requiring federal approval before any release. The administration planned to finalize that framework the same week, which makes the Washington demo as much a scientific event as an exercise in lobbying.
The contrast with the rest of the lineup is sharp. While Astra targets fundamental research, OpenAI pushes its everyday models toward daily use, like ChatGPT Voice now controlling the whole desktop. Two fronts, one logic: hold both the scientific peak and the consumer volume.
The usual caution applies. Ten proofs formalized and checkable in Lean are a serious signal, but a hand-picked demonstration says nothing about the model’s reliability on a problem chosen at random. The real test comes when outside researchers can run Astra on their own questions.
The regulatory calendar will weigh heavily. If the federal framework tightens, Astra could stay in the display case for months while approval catches up. The model that cracks decade-old problems will first have to wait at a brand-new administrative counter.
For rival labs, the pressure is now qualitative, not just numeric. Beating an exam benchmark stops being the headline when a competitor claims original results in fields where no benchmark exists. The next move everyone will watch for is whether Astra’s approach reproduces outside a curated set, because a capability that only works on hand-picked problems is a demo, and one that works on a stranger’s question is a product.
Follow the story on Horizon.


