Astra lands with a 99.9% score on ARC-AGI-3 and a bill 2.5 times heavier than GPT-5.6 Sol. This GPT-6 Astra test pulls together the published evaluations, the real rates and the measurement conditions, to work out where the premium holds and where it does not.
Key Takeaways
- The 99.9% on ARC-AGI-3 comes from a custom harness; the benchmark’s own standard harness returns 62.7%.
- Astra dominates machine use and cybersecurity, yet trails Fable 5.1 on the aggregate index and on tool-assisted reasoning.
- Caching decides the final invoice: 1 dollar per million on reads, 12.50 on writes.
Have an AI Sum Up This Article
ChatGPTWhat the Harness Makes a Score Say
The first move in this GPT-6 Astra test was to separate the figure from its method. The 99.9% on ARC-AGI-3 was produced with a provider adapter written for the occasion, not with the benchmark’s default scaffolding.
Run through the standard harness, the same model on the same test drops to 62.7%. A thirty-seven point spread does not measure capability, it measures the tooling wrapped around it.
The caveat does not cover everything, thankfully. The model had already delivered a result no scaffolding can manufacture, by closing ten mathematics problems that had stayed open, and that kind of outcome survives any argument about method.
That distinction stops being methodological trivia the moment a team plans a migration. The number at the top of the product page describes a ceiling reachable with bespoke code, not the behaviour of the model wired into an existing pipeline.
The same problem recurs in another shape. No cross-vendor coding benchmark run inside a single harness currently includes Astra, which makes any head-to-head coding table hard to take at face value.
The one standout independent measurement in this segment belongs to a rival anyway. Claude Opus 5 posts 97.0% on SWE-bench Verified at Vals AI under a minimal bash-only harness, and nobody has yet run Astra through the same setup.
The cyber results carry a caveat of the same order. The 100% on ExploitBench and the 42.4% on ExploitGym were recorded without production safeguards, so on a model nobody will ever run in that state.
The practical consequence is easy to state. Until those conditions are reproduced under real deployment settings, the cybersecurity column on the spec sheet describes a maximum potential rather than a shipped capability, and no buying decision should rest on it.
Where Astra Genuinely Pulls Away
Set the questionable scores aside and a solid core remains. Machine use is the area where the model gains most clearly on the previous generation.
OSWorld 2.0 rises to 72.6% against 65.7% for GPT-5.6 Sol, with tasks completing 47% faster. ScreenSpot-Pro moves from 76.9% to 92.7%, close to sixteen points on locating elements on screen.
Inside an agent loop, the speed gain matters more than the score gain. A task that finishes twice as fast burns fewer turns, therefore fewer output tokens, which offsets part of the headline premium.
On structured professional work, the margin turns decisive. AutomationBench reads 41.4% against 31.4% for Fable 5.1 and 18.1% for Sol, while BenchCAD shows 95.9% against 84.3%.
Long-context reading holds up too. Accuracy stays at 100% between 256,000 and 512,000 tokens, then 96.3% beyond, across a window that accepts 922,000 tokens of input.
Cyber remains the most spectacular field despite the method caveats. SRE-Bench climbs from 55.9% to 88.0%, and the model surfaced two unknown zero-day flaws during its own evaluation, which explains the Critical rating anticipated back in August.
On code, by contrast, the margins narrow until they vanish. Terminal-Bench 4.0 gives 57.7% against 55.8% for Fable 5.1, FrontierCode 1.1 Main gives 53.3 against 53.5 for Fable 5, and FrontierCode Extended puts Fable 5 outright ahead at 64.9% against 64.5%.
DeepSWE v1.1 completes the picture. Astra records 74.1% there, ahead of Sol at 72.7%, but behind Muse Spark 1.3 at 75.4% for a rate that is not remotely comparable.
The aggregate index confirms that reading. Artificial Analysis places Astra at 61.2, level with the previous generation, where Fable 5.1 holds 65.7.
More articles on Horizon
- GPT-6 Astra Launches as OpenAI Claims AGI Era
- OpenAI Makes Astra’s Reasoning Harder to Follow
- Gemini 3.8 Flash Trails Claude Opus 5 on Coding
Our Verdict, Workload by Workload
The grid is easy to memorise. Ten dollars per million on input, fifty on output, one dollar for cached input and twelve dollars fifty for cache writes.
That last figure decides the real invoice. A workload with a stable prefix absorbs most of the premium, while a workload rewriting its context on every call stacks up cache writes billed at twelve times the read rate.
Fast mode answers to the same logic. It multiplies the total by two and a half for a matching speed factor, which makes it a latency trade, never a saving.
The available tooling weighs into the decision, and it is comprehensive. The model takes web search, file search, a code interpreter, a hosted shell, patch application, computer use and the MCP protocol.
For an agent driving an interface or an office pipeline, adoption needs no argument. It is the one field where the measured gaps are wide, reproduced across several benchmarks and consistent with the stated speed gain.
For a team writing code all day, our verdict is to wait. One or two points collected in mismatched harnesses do not pay for a rate two and a half times the previous generation, especially when a rival model edges ahead on DeepSWE.
The context window deserves its own line in that arbitration. A 922,000-token input ceiling holding 96.3% accuracy at the far end removes the chunking layer that most document pipelines still carry, and maintaining that layer has a cost of its own.
The April 30, 2026 knowledge cutoff sets a boundary worth noting too. Anything published after that date reaches the model through search or retrieval, never from its own weights, which matters for fast-moving technical documentation.
For vulnerability research, access stays filtered regardless. The sharpest offensive capabilities go to testers first, which pushes the question out to a rollout whose calendar is not public.
One point this test cannot settle today, for lack of distance. Launch data is too thin to confirm that efficiency gains cover the pricing premium across weeks of real load, and the general intelligence claim attached to the launch changes nothing in that arithmetic.
We will run the numbers again once a shared harness has put Astra and its rivals through identical conditions. Until then, the product page describes a ceiling, and the invoice describes the day job.
Follow the story on Horizon.


