GPT-6 Astra Launches as OpenAI Claims AGI Era

GPT-6 Astra shown on twin control room dials giving two opposite verdicts

OpenAI put GPT-6 Astra into limited preview on September 3, priced at 10 dollars per million input tokens and 50 per million output. The model comes out of the largest training run the lab has ever attempted, and its unveiling carried a line nobody there had risked before.

Key Takeaways

  • Astra posts 99.9% on ARC-AGI-3, then drops to 62.7% once the benchmark’s standard harness is used instead.
  • It hits 100% on ExploitBench and becomes the first model rated at OpenAI’s Critical cybersecurity threshold.
  • On the Artificial Analysis index, Claude Fable 5.1 still leads at 65.7 against 61.2.

Have an AI Sum Up This Article

ChatGPT

A Hundred Thousand GPUs Behind the Lab’s Largest Run

GPT-6 Astra shipped on September 3 as a limited preview restricted to organisations inside the Daybreak program. The rollout to Plus, Pro, Business and Enterprise accounts follows within days, alongside the API, AWS Bedrock and Azure.

Two variants are in circulation. The standard build carries the gpt-6-astra label in the API, while Astra Pro stays tied to Pro, Business and Enterprise plans.

Training drew on more than 100,000 GPUs at the Stargate site in Texas. OpenAI frames it as its biggest campaign to date, which sits oddly against August, when the lab put its heaviest training run on hold.

The context window lands at 1,050,000 tokens, with a 922,000 ceiling on input and 128,000 tokens of output. The knowledge cutoff is set at April 30, 2026.

Greg Brockman summed the launch up in four words, “Welcome to the AGI era”, arguing the model may qualify as artificial general intelligence or mark its threshold. Inside OpenAI, that register had been reserved for distant roadmaps until now.

The wording carries weight for the lab. Where exactly a system crosses that line has already occupied a courtroom, when an expert witness warned about the arms race such a crossing would trigger.

The most telling precedent is mathematical. The model had already drawn attention by closing ten open problems nobody had settled before it, a month ahead of general availability.


GPT-6 Astra

The Scores That Carry the Claim, and the Ones That Undercut It

The headline figure is 99.9% on ARC-AGI-3. It falls to 62.7% once the evaluation runs through the benchmark’s standard harness, the one used to line models up against each other.

A thirty-seven point gap between two readings of the same test restates a plain rule. The score owes as much to the scaffolding driving the model as to the model itself, and the two do not compare.

Elsewhere the results hold without qualification. FrontierMath Tier 4 reaches 97.6%, GPQA Diamond 96%, and long-context reading stays at 100% accuracy between 256,000 and 512,000 tokens, then 96.3% beyond that.

On machine use, Astra pulls clearly ahead of GPT-5.6 Sol. OSWorld 2.0 moves from 65.7% to 72.6% with tasks running 47% faster, and ScreenSpot-Pro climbs from 76.9% to 92.7%.

Structured professional work follows the same pattern. AutomationBench reads 41.4% against 31.4% for Fable 5.1 and 18.1% for Sol, while BenchCAD lands at 95.9% against 84.3%, two gaps wide enough to survive a change of harness.

Cybersecurity is where the margin turns spectacular. ExploitBench reads 100%, ExploitGym 42.4% against 30.3%, and the model surfaced two unknown zero-day flaws during its own evaluation.

Those numbers come with a footnote. They were produced without production safeguards in place, which explains why the Critical cybersecurity threshold was already anticipated back in August.

Access to the sharpest offensive capabilities stays filtered as a result. Testers come first, ahead of any widening through the cyberdefence platform the lab opened in the spring.

The counterweight sits in the aggregate ranking. Astra scores 61.2 on the Artificial Analysis index where Claude Fable 5.1 holds 65.7, and that distance is awkward for a model sold as a generational break.

Task-level detail sharpens the point. Fable 5.1 wins Humanity’s Last Exam with tools at 65.0% against 57.2%, and Muse Spark 1.3 edges past on DeepSWE v1.1 with 75.4% against 74.1%.


More articles on Horizon


What the Bill Changes for Teams Planning a Move

The rate matches Anthropic’s top tier to the dollar, 10 per million on input and 50 on output. Two premium grids converging that precisely is not a market accident.

Measured against the previous generation, the bill bites harder. GPT-6 Astra costs 2.5 times GPT-5.6 Sol, and fast mode multiplies the total by another two and a half for a matching speed gain.

Caching reshapes how that bill reads. Cached input drops to 1 dollar per million, but cache writes run at 12.50, which punishes workloads whose context shifts on every call.

For a product team, the decision therefore turns on context stability rather than on ranking. An agent loop replaying the same prefix absorbs the premium, while a chain that rebuilds its prompt each turn pays full freight.

The knowledge cutoff adds a second variable to that calculation. Set at April 30, 2026, it means anything more recent has to be fed in through search or retrieval, which lengthens prompts and pushes input volume back up on exactly the workloads the premium is meant to serve.

On the competitive side, the launch squeezes two fronts at once. Anthropic loses its pricing edge at the top, and Meta now has to defend a cost advantage with a model that beats Astra on code for a fraction of the spend.

The auditability question stays open, and no score table settles it. Astra leans on recurrent depth, pushing part of the computation outside the readable trace, a choice we unpacked while watching the model’s reasoning become harder to follow.

The coming weeks will show which of the two signals holds. A generational claim rests on scores gathered under favourable conditions, while the aggregate index and the invoice tell a far more ordinary story.

Follow the story on Horizon.

Comments

No comments yet. Why don’t you start the discussion?

    Leave a Reply

    Your email address will not be published. Required fields are marked *