Qwen 3.8 Max Catches Claude Opus but Hallucinates More

Qwen 3.8 Max sharing a podium step with Claude Opus while a scale tips toward the invoice side

Qwen 3.8 Max, Alibaba’s latest model, pulls level with Claude Opus 4.8 on the reference intelligence index, scoring 56 against 57 for Kimi K3. Its cost per task runs a third above the rival Chinese model, and its hallucination rate has almost doubled.

Key Takeaways

  • Qwen 3.8 Max scores 56 on the Artificial Analysis index, level with Claude Opus 4.8 and one point behind Kimi K3.
  • One task costs $1.14 on Qwen 3.8 Max against $0.86 on Kimi K3 and $0.57 on GLM-5.2.
  • The hallucination rate climbs from 23% to 40% and the general knowledge index drops 10 points.

Have an AI Sum Up This Article

ChatGPT

One point separates the top of the intelligence board

On the Artificial Analysis intelligence index, the hierarchy fits in three numbers. Kimi K3 leads on 57 points, Qwen 3.8 Max and Claude Opus 4.8 sit on 56, and GLM-5.2 trails on 51.

A single point between first and second is statistical noise. What matters is that a general-purpose Chinese model now sits level with a US frontier model on a composite index, with no asterisk attached.

The model topping that board is not a closed product either. Moonshot went the open route with the release of Kimi K3’s weights, 1.4TB handed over for free, which allows in-house hosting and rewrites the cost equation entirely.

The ranking reshuffles on work-related tasks. On GDPval-AA, Claude Opus 5 leads comfortably with 1,852 Elo points, Qwen 3.8 Max follows on 1,739 and Kimi K3 slips to 1,685.

Read together, those two boards say something useful. Qwen moves ahead of Kimi as soon as you measure real professional tasks, while staying behind on the general index. The two models are not strong in the same places, and a buyer picking on a single headline number will end up with the wrong one for their workload.

Release context matters too. Alibaba opened the sequence in late July, when the group shipped Qwen 3.8, its biggest AI model to date. The Max version lands a few weeks later on a premium footing, aimed squarely at buyers who were previously locked into US frontier pricing.


Qwen 3.8 Max

Cost per task flips the trade-off

Qwen 3.8 Max prices cleanly: $2.00 per million input tokens, $6.00 on output, and $0.25 when a request lands on a cache hit. Nothing unusual at this capability tier.

List price says very little on its own. What a team running this at scale cares about is the cost of a task carried through to completion, and there the gap opens: $1.14 on Qwen 3.8 Max, $0.86 on Kimi K3, $0.57 on GLM-5.2.

The reason sits in how the model works. Qwen 3.8 Max burns 64 steps per task where the alternative needs 14, and input token volume grows fifteenfold because the full conversation history gets resent on every turn.

That implementation detail is expensive at volume. Across several thousand tasks a day, a third on the invoice weighs far more than one point on a composite index nobody will ever see in production. Finance teams read invoices, not leaderboards, and that is the conversation an integration lead ends up having internally.

The maths turns friendlier wherever the cache does real work. At $0.25 per million cached input tokens, a workload built on highly repetitive prompts absorbs part of that step overhead.

The usage profile has to cooperate, though. A support assistant with a stable system prompt and short variations gets the full benefit of caching. An agent exploring a different code repository on every run will barely touch it.

The systematic history resend deserves a note of its own. That is exactly the kind of behaviour a vendor fixes within weeks through better context handling. The cost gap measured today is nothing like permanent.


More articles on Horizon


Alibaba moves fast, with a visible regression attached

The downside of the upgrade shows up in the reliability numbers. The hallucination rate moves from 23% to 40%, with accuracy sitting around 31%.

Two more indicators slip in the same direction. AA-LCR loses 2 points, and AA-Omniscience, which measures breadth of general knowledge, falls 10 points against the previous generation of the model.

Losing general knowledge while gaining on task execution points at where the training effort went. Broad recall costs capacity that can be spent elsewhere, and the numbers suggest Alibaba spent it on agentic behaviour rather than on knowing things.

A model that gains on agentic reasoning while losing on factual accuracy is not stumbling, it is making a trade. It suits tooled pipelines where every output gets checked, and suits far less any workflow where the answer goes straight to a user.

Deployment mode changes that reading too. A proprietary model billed per token has to hold its reliability promise on every call, because every retry costs money. A self-hosted model absorbs retries differently, which makes imperfect accuracy easier to live with.

Hardware remains the sector’s real ceiling. Moonshot had to halt Kimi sales in July for lack of available GPUs, a reminder that a pricing advantage only counts when compute capacity keeps up behind it.

Alibaba’s publication rhythm is worth pausing on. Two major models in under a month, with a positioning jump between them, leaves integration teams very little room to qualify one version properly before the next one lands.

For US labs the signal cuts both ways. Raw parity on the general index has been reached, yet the gap holds on professional tasks, where Claude Opus 5 keeps more than a hundred Elo points of headroom. That is the ground the next response will be fought on.

Follow the story on Horizon.

Comments

No comments yet. Why don’t you start the discussion?

    Leave a Reply

    Your email address will not be published. Required fields are marked *