Zhipu has pulled the mask off Ox Alpha: the stealth model that spent a week at the top of the usage charts is GLM-5.3-Flash, and it was served from a cluster of 100,000 domestically produced chips. The company’s stock closed more than 12 percent higher in Hong Kong, and the weights are now public under an MIT licence.
Key Takeaways
- GLM-5.3-Flash carries 320 billion parameters with 18 billion active, and becomes the first natively multimodal model in the GLM-5 series.
- The serving cluster runs on 100,000 Chinese chips, and the model absorbed 62 trillion tokens before its formal release.
- Pricing lands at $0.15 per million input tokens and $0.50 on output, halved until September 9.
Have an AI Sum Up This Article
ChatGPTThe mask comes off and the stock jumps
We had been tracking the question for a week. Ox Alpha had landed on OpenRouter with no vendor name and no bill attached, and testers split between two theories, a Chinese lab close to the GLM family or Microsoft’s in-house model programme.
The first theory won. Zhipu confirmed that the model tested under a pseudonym belonged to its GLM series, under the name GLM-5.3-Flash, and released the weights the same day.
Markets moved within hours. The stock closed up more than 12 percent at HK$1,160 on Thursday, a reaction that rewards the staging of the launch as much as the model itself.
The usage numbers built up during the anonymous run explain the enthusiasm. The model had processed 62 trillion tokens before its formal release, 11 trillion of them on OpenRouter within three days, the biggest launch the platform has recorded.
The method behind the anonymity is worth naming. A stealth release buys a lab weeks of large-scale feedback without putting its brand behind a result it cannot yet guarantee, and it turns thousands of paying-nothing testers into an evaluation team.
The breakdown of those volumes says more. Coding alone accounts for 10.3 trillion tokens, or 31 percent of the platform’s weekly throughput, which pushed the model to first place among coding systems in use.
The lab has form on this ground. Zhipu already shook the ranking when it shipped a free rival to Claude with ZCode, though the scale reached this time changes the nature of the exercise.
One hundred thousand domestic chips, and what that proves
The heaviest technical point is not about the model but about the hardware serving it. The stealth run went through a cluster of 100,000 chips produced in China, at a scale and under a load nobody had documented publicly until now.
The demonstration is political as much as technical. Serving tens of trillions of tokens without American silicon answers the question export restrictions have been posing for two years, which is whether China can hold inference at scale.
It also ran in the open. For a week thousands of developers wired their pipelines into that cluster without knowing what sat behind it, and none of them reported the service degradation that would have exposed second-tier infrastructure.
GLM-5.3-Flash shows 320 billion parameters with 18 billion active per token, a ratio designed for cheap inference. It is also the first in the GLM-5 series to be natively multimodal, taking text, images, video and visual documents as input.
Part of the gain sits in the architecture. Zhipu combines sparse and linear attention for the first time in the series, and adds constrained hyper-connections, a block labelled mHC meant to improve scaling efficiency.
The weights carry no strings. The official card published by Z.ai puts the model under an MIT licence and places it near Claude Opus 4.8 on coding and agentic work.
The sequence mirrors the other Chinese labs. Moonshot opened the path when Kimi K3 published 1.4 TB of open weights, and systematic weight releases have become the shared commercial argument of the entire local sector.
More articles on Horizon
- Codex Prepares a Mode That Never Stops Working
- Qwen3.8-Flash-Next Cuts AI Prices Twelvefold
- Anthropic IPO Targets a $2 Trillion Valuation
A price aimed straight at American APIs
The published price finishes the argument. Input runs at $0.15 per million tokens, output at $0.50, and cached input drops to $0.03, a level where cost stops being a selection criterion.
A launch promotion halves those numbers until September 9. For two weeks a team can therefore evaluate a frontier-class model at $0.075 per million input tokens, roughly what a small previous-generation model costs.
For product teams the evaluation window is open, dated and easy to justify internally. The MIT weights also allow an internal deployment, which answers the data retention objection raised during the model’s anonymous phase.
The shift from free to paid is the part worth watching. During the preview every token was served at no charge, so the meaningful number this week is not the discount but the first invoice, and teams that built a habit on a free endpoint now have to price it into their agent runs.
The hardware question looks different from the buyer’s seat. A model proven on Chinese chips remains a model you can run elsewhere, but the sovereignty argument it carries will land first with buyers trying to escape dependence on a single GPU supplier.
On the competitive side the pressure hits the neighbours first. Alibaba had just dropped to $0.16 per million input tokens with Qwen3.8-Flash-Next, and finds itself undercut on its own ground less than twenty-four hours after the announcement.
American labs take the hit indirectly. An open model claiming the neighbourhood of Opus 4.8 on code, at a tenth of premium subscription pricing, mechanically shrinks the space left to justify closed offerings on development work.
The local release cadence blurs the picture further. Between Alibaba shipping Qwen 3.8 this summer, Moonshot’s releases and this drop from Zhipu, three Chinese labs have delivered competitive open models in under two months.
Our read fits in one line. The real signal this week is not the model’s score but the cluster that served it, because a software gap closes in a matter of months while a hardware gap is counted in years of fab capacity.
Follow the story on Horizon.


