xAI shipped Grok 4.5 on Wednesday July 8, 2026, its first model purpose-built for coding and agentic workflows. The model rivals Claude Opus 4.7 on announced benchmarks, at $2 per million input tokens and $6 per million output tokens, and burns roughly four times fewer output tokens on hard tasks. The launch lands one day before OpenAI’s GPT-5.6 worldwide rollout and reshuffles the pro coding segment.
Key Takeaways
- xAI shipped Grok 4.5 on July 8, 2026, its first model purpose-built for coding and agentic tasks.
- Pricing at $2 input and $6 output per million tokens, more than 60% cheaper than Opus 4.8 and GPT-5.5.
- Mixed benchmarks: Grok 4.5 beats Opus 4.8 on DeepSWE 1.0 and Terminal-Bench 2.1, loses on DeepSWE 1.1 and SWE-Bench Pro. Score 54, ranked 4th on the Intelligence Index.
Have an AI Sum Up This Article
ChatGPTA model built for coding and agents from day one
Grok 4.5 is the first xAI model explicitly built for coding and agentic workflows. Prior Grok releases targeted a general-purpose chatbot tuned for the X ecosystem. This time, Musk announced a pivot: stay competitive on the long-horizon, iterative, tool-calling coding tasks where the enterprise stack now lives.
In the launch note, Musk framed Grok 4.5 as an “Opus-class” model, then narrowed the comparison: “roughly comparable to Opus 4.7, but much faster”. The line aims squarely at Anthropic, whose Claude Opus has held the reference slot on high-end coding for a year, next to a stack tuned for autonomous agents.
The Artificial Analysis Intelligence Index ranking confirms a solid slot. Grok 4.5 lands at 4th place with a score of 54, ahead of every Gemini and every open-weight model that ships today. That result reinserts xAI into the top tier a year after the brand slipped behind Anthropic and Google on the pro segment.
The agentic turn deserves a separate line. Multi-tool orchestration has become the real applicative differentiator, as we unpacked in our coverage of Meta’s AI agents running slower than Zuckerberg expected. Musk pushed Grok 4.5 on the same ground, leaning on latency and per-call cost rather than raw quality on an isolated single-turn answer.
Another shift worth logging: Grok 4.5 arrives without a new consumer-facing version of the brand. Previous xAI efforts poured energy into muscling the conversational experience on X. This time, the stated target is the API and pro integrations, with direct availability on the major hyperscalers for heavy coding workloads.
Mixed benchmark verdict, a real punch on tokens
The scoreboard reads mixed on raw performance. Grok 4.5 outperforms Opus 4.8 on DeepSWE 1.0 with 62% resolution against 55.75%, and on Terminal-Bench 2.1. But Opus 4.8 stays ahead on DeepSWE 1.1 at 59% against Grok’s 53%, and on SWE-Bench Pro with a 69.2% resolve rate against 64.7%.
The direct read is simple: on the newest and hardest benchmarks, Anthropic holds a lead. xAI does not claim the top slot on raw resolution rate. The brand built its case on a second axis worth examining closely: the amount of tokens burned to solve a given task.
On the Intelligence Index, Grok 4.5 averages 14,000 output tokens per task, against 67,020 for Opus 4.8 on the same set. On SWE-Bench Pro, the ratio holds: xAI estimates the average consumption at around 15,954 tokens per task, when Opus 4.8 in maximum-effort mode still burns the same 67,020 tokens.
What that changes on the bill: for a comparable coding task, Anthropic billing can end up roughly four times higher than xAI at similar resolution rates. The reasoning tracks the pricing move Anthropic pushed last week and documented in the Fable 5 switch from subscription to pure token billing, where the per-call load became the central adjustment variable for integration teams.
Token efficiency reframes the whole conversation. Classic benchmarks measure the resolution rate, not the unit cost of that resolution. Musk pointed to the gap, and to how the absolute Opus score gets paid for as soon as one stacks several thousand agentic requests a day inside a production pipeline.
Grok 4.5 pricing lands at $2 per million input tokens and $6 per million output tokens. At similar resolution rates, the saving on a billing envelope of several billion tokens a month becomes the central argument for a CIO arbitrating between top-tier performance and coding stack running costs.
Also on Horizon:
- GPT-5.6 Test: We Rate Sol, Terra and Luna
- GPT-Live: ChatGPT Now Listens and Talks at Once
- GPT-5.6 Launches Worldwide From OpenAI Today
Impact on users and rivals: the deck reshuffles
On the user side, the window opens for teams shipping heavy coding workloads. Engineering shops pipelining tests, refactors, docstring generation, or platforms delivering tailored assisted code: these workloads burn tokens by the ton. A swap from Opus 4.8 to Grok 4.5 at comparable quality on DeepSWE 1.0 lightens the monthly bill in a way that shows up on a P&L slide. We put the promise to the test, our full Grok 4.5 test.
The choice becomes an actual product call. A team betting on high resolution on rare but critical tasks stays on Opus 4.8 for the cases where the SWE-Bench Pro margin pays back. A team absorbing a dense flow of automated coding iterations can switch to Grok 4.5 with no major loss on the task average and a cash win at the end of the month.
On the rival side, Anthropic keeps the premium slot on absolute resolution but loses the budget-friendly differentiator it had held between GPT-5.5 and open-weight models. OpenAI sees its GPT-5.5 mid-tier inside Codex directly challenged: Musk claims a per-task cost roughly half of GPT-5.5 in the same Codex. Google takes one more step back, with the Gemini line pushed behind Grok 4.5 on the Intelligence Index.
The calendar is tight. Grok 4.5 shipped on July 8, one day before the GPT-5.6 worldwide rollout unpacked in our earlier coverage of the OpenAI GPT-5.6 global rollout. The Sol, Terra and Luna wave arrives at a moment when xAI has already set the new budget standard. The direct comparison between GPT-5.6 Terra at $2.50 input and $15 output and Grok 4.5 at $2 and $6, at comparable declared performance, will weigh on the next API migrations.
The next rival move should surface within weeks. Anthropic will have to choose between defending its premium margin on Opus 4.8 or aligning an intermediate model calibrated on the same cost logic. OpenAI will probably lean on the workspace agents suite freshly announced with GPT-5.6 rather than enter a tariff war on Luna. The mid-tier LLM battlefield becomes the real front of Q3 2026.
Follow the story on Horizon.


