SpaceXAI shipped Grok 4.6 on Wednesday and the model landed level with GPT-5.6 Sol Max on the industry’s reference index. Its price card sits more than 60 % below what Anthropic and OpenAI charge for the same tier of work. Musk’s lab just turned model selection into a billing decision.
Key Takeaways
- Grok 4.6 scores 61 on the Artificial Analysis index, tying GPT-5.6 Sol Max
- Input runs at $2 per million tokens against $5 for Claude Opus 5 and GPT-5.6 Sol
- The base model never grew, every gain came out of post-training
Have an AI Sum Up This Article
ChatGPTThe 61 That Puts SpaceXAI Level With Sol Max
Grok 4.6 went live on Wednesday and its aggregate intelligence result fits in two digits: 61. That number clears Kimi K3, Moonshot’s open-weight heavyweight, and lands exactly on GPT-5.6 Sol Max, a tie no xAI release had managed before.
The full board, kept by the independent shop that compiles the frontier intelligence index, still shows Claude Opus 5 on top with Fable 5 right behind. Anthropic holds both leading spots, and the gap down to third just closed.
On GDPval-AA v2, which tries to score genuine knowledge work done on a computer, Grok 4.6 sits second with an Elo of 1,753. Claude Opus 5 is the only model ahead of it there. For buyers shopping execution rather than leaderboard points, that second place carries more weight than the aggregate score.
The engineering detail is the sharpest part of this launch. SpaceXAI held the base model constant and routed every capability gain through post-training. The previous generation had already tried that play with a 4.5 release priced at a third of what Claude Opus was charging, without reaching the top three.
A year ago the same jump would have demanded a fresh base and a full training run, along with the months of compute booking that come with it. A post-training pass closing fifteen leaderboard places says something concrete about where the race actually stands: raw scale is no longer the only lever on the table.
Two Dollars per Million Tokens Is the Real Story
Below 200,000 prompt tokens, Grok 4.6 bills $2 per million on input, 50 cents on cached input and $6 per million on output. Cross that line and the whole card doubles to $4, $1 and $12.
Set against the competition, the spread is blunt. Claude Opus 5 asks $5 in and $25 out, GPT-5.6 Sol asks $5 and $30. At comparable stated quality, the distance runs past 60 % on output, which happens to be the line that balloons on agentic workloads.
That 200,000-token threshold deserves a second look, because the sticker price and the invoice rarely match. An agent session stacking a repository, runtime logs and several rounds of tests crosses the line without warning, and flips onto the upper card.
Even at $4 in and $12 out, the model still undercuts the entry rate of both direct rivals. On long contexts the gap narrows, and it never flips.
For product teams, that moves the whole calculation. An agent grinding for hours through an unfamiliar codebase burns tens of millions of tokens a week, and that cost line now decides whether a use case ships or stays a demo behind the login.
SpaceXAI also put the model everywhere at once. Grok Build opens it from the $30 a month SuperGrok plan, Cursor carries it on the editor side, and OpenRouter, Vercel and Cloudflare relay it on infrastructure. The same outfit is the one building a consumer handset Musk publicly denied back in July.
Distribution is doing real work here. A model that shows up inside the editor a developer already has open, and behind a router a platform team already pays for, gets evaluated far more often than one that requires a new account and a procurement conversation.
The lab frames the product around long-horizon agents, coding and heavier visual projects. It leans on the model’s ability to stay on task across many steps, work through a codebase it has never seen, and check its own output before moving forward.
More articles on Horizon
- Riot Platforms Rents Its Data Center to Anthropic
- ChatGPT Business Adds a $125 Premium Seat
- Claude Watermark Now Covers All Generated Text
What a Halved Bill Forces on Every Other Lab
Anthropic keeps the quality crown with two models above Grok 4.6. The commercial story is what cracks: defending a two-and-a-half times premium gets awkward once the third-place model clears the same workloads.
OpenAI sits in the tighter spot. GPT-5.6 Sol Max has now been caught on score and beaten on price, which leaves three moves: cut the card, differentiate through the tooling wrapped around the model, or pull the next release forward.
Switching costs work in SpaceXAI’s favour too. Pointing an agent at a different model from Cursor or through an API router is an afternoon of work with nothing signed, which makes the comparison easy to run for a team still hesitating.
There is a second-order effect worth watching on the enterprise side. Procurement teams that spent the last year standardising on a single frontier vendor now have a documented reason to reopen the file, and every reopened file is an opening for whoever prices lowest at acceptable quality.
The message reads the same from the Chinese side. Kimi K3 built its reputation on performance per dollar, and watching a US model pass it on the index while staying aggressive on billing strips Moonshot of its main adoption argument outside China.
One variable stays off the leaderboard. Musk’s lab still carries open questions about how it runs internally, including the June lawsuit from a safety engineer who says he was pushed out after raising alarms. Engineering leaders signing a multi-year contract weigh that alongside the rate card.
Two responses look plausible over the coming weeks: a price adjustment from OpenAI on the Sol tier, or a product answer from Anthropic that drags the argument away from tokens and onto agent tooling. Standing still is what nobody can afford now.
Follow the story on Horizon.


