Artificial Analysis Rebuilds Its AI Ranking

Artificial Analysis swaps the exam mid-session as robot candidates freeze in disbelief

GPT-6 Astra landed on 61 points on the Intelligence Index, the exact score of the model it replaced, and the backlash was immediate. Artificial Analysis has now shipped version 4.2 of that ranking, which hands OpenAI’s model four points back. Two benchmarks come in, one goes out, and private test data now carries 40 percent of the weighting.

Key Takeaways

  • Astra scored 61 at maximum reasoning effort, matching GPT-5.6 Sol point for point
  • Version 4.2 of the index gives it a four-point gain over its predecessor
  • Claude Fable 5.1 still leads the ranking, with Astra second and Meta third

Have an AI Sum Up This Article

ChatGPT

Sixty-one points for a model sold as a break

The number aged badly within hours. Pushed to maximum reasoning effort, Astra came out at 61 points on the Artificial Analysis Intelligence Index, which is exactly what GPT-5.6 Sol scores, the model it was meant to supersede.

The gap with every other measurement was impossible to miss. Other test suites placed Astra well above Sol, a contrast made louder by the fact that OpenAI framed the launch as the start of the AGI era.

The criticism landed on the instrument, not the model. A ranking that returns the same figure for two consecutive generations says nothing about real progress, it mostly says its tests no longer capture what changed between them. That is a familiar failure mode for any benchmark that survives more than a couple of model cycles.

Artificial Analysis had in fact measured a gain elsewhere. On its coding agent index, Astra pulls level with Fable 5 at a lower cost, a result that sat awkwardly beside the flat score on the general index.

It also clocked an efficiency improvement. Astra burns fewer tokens than Sol for comparable performance, an edge cancelled out by higher pricing, and the outfit walks through all of those readings in its measurement report on GPT-6 Astra.

That pricing point matches what we flagged at launch. The generational jump comes with a bill, and our own test of the model already weighed what Astra’s price actually buys.


Artificial Analysis

AA-Briefcase and GDP.pdf in, GPQA-Diamond out

Version 4.2 changes what the exam contains. Two benchmarks arrive, AA-Briefcase covering real-world knowledge work, and GDP.pdf supplied by Surge AI, which grades analysis of PDF documents.

One veteran test leaves in the same pass. GPQA-Diamond is dropped because models now solve it, a textbook saturation case where an exam stops separating candidates once they all pass. Retiring it says as much about the pace of the field as any single model release does.

The structural change is about the data itself. Private test sets now carry 40 percent of the overall weighting, a share designed to make targeted optimisation far harder for labs that would otherwise know the public exams by heart.

The recalculation produces a clean result. Astra now shows a four-point lead over its predecessor, which puts the generation back on a progress curve without handing it the top spot.

Anthropic keeps that spot. Claude Fable 5.1 stays in front, Astra takes second and Meta rounds out the top three, a hierarchy consistent with what our test of Fable 5.1 observed on cache pricing and code quality.


More articles on Horizon


A ranking that became a piece of the market

For teams choosing a model, the useful lesson is methodological. An index that changes composition mid-year rules out any direct comparison between two scores weeks apart, and the version number becomes as load-bearing as the score itself.

The practical consequence states itself. An infrastructure decision resting on a general ranking now needs a second layer, an internal test on the team’s actual tasks, or it ends up tracking a third party’s revisions instead of its own requirements. Building that internal test costs a few days of work, which is cheap against the price of migrating a production stack twice in a quarter.

On the lab side, the episode sets an uncomfortable precedent. An independent ranking that revises itself after public pushback invites questions about its authority, and that authority is precisely what makes it worth anything to buyers. Every future update will now be read against the suspicion that a loud enough complaint can move the scoring.

The private data share is the answer to that fragility. By keeping part of the exam invisible to labs, Artificial Analysis shields itself from the heaviest charge facing every public ranking, that it measures training on the test rather than general capability.

This deliberate opacity arrives in an already opaque context. External measurement matters more as model transparency shrinks, given that Astra’s reasoning has become harder to follow for anyone trying to audit it.

The next revision will test the outfit itself. If another contested model triggers a fresh method update, the question stops being about a score and becomes whether a private ranking can stay the referee of a market now worth tens of billions of dollars a year.

Follow the story on Horizon.

Comments

No comments yet. Why don’t you start the discussion?

    Leave a Reply

    Your email address will not be published. Required fields are marked *