Moonshot has suspended new subscriptions to Kimi K3 after demand pushed its compute capacity to the edge in forty-eight hours. Existing subscribers are untouched, and new slots will reopen in batches.
Key Takeaways
- Moonshot stopped selling new subscriptions, capacity maxed out in 48 hours.
- The two affected plans are Kimi Membership and Kimi Code Membership.
- The model runs 2.8 trillion parameters and leads on frontend code.
Have an AI Sum Up This Article
ChatGPTForty-eight hours to fill every available machine
Moonshot announced the freeze by saying demand had pushed the company close to the limits of its current capacity. The window is the real story here. Forty-eight hours.
A lab that shuts its commercial door two days after a launch is not running a marketing play. It has run out of free machines to serve one more customer without degrading the ones already paying.
The model behind the rush is not a light one. Kimi K3 carries 2.8 trillion parameters, and we walked through last week how this open model went up against the strongest American systems.
The published scores explain the traffic. On Code Arena Frontend, K3 posts 1,679 points ahead of Fable 5 at 1,631 and GPT-5.6 Sol at 1,618, a first for a Chinese model on that board.
The picture flips elsewhere. On FrontierMath Tier 4 the model stalls around 39 percent while the OpenAI and Anthropic systems land near 90 percent, a gap the developers who came for the coding clearly shrugged off.
That asymmetry goes some way to explaining the rush. Frontend code is a high-volume, immediately measurable use case, while advanced mathematical reasoning stays a niche workload for most product teams.
The sequence also says something about the initial sizing. Moonshot launched a model capable of topping an international leaderboard on a compute reserve that did not survive two days of global curiosity.
Two plans pulled from sale, current members left alone
The freeze covers two restructured offers. Kimi Membership spans the web, the app and the work features, while Kimi Code Membership targets programming workflows.
Protecting existing subscribers fits the shape of the problem. When the constraint is physical, taking on new customers means splitting the same silicon across more sessions, which slows everyone down at once.
Users are already reporting execution speeds well below the American alternatives. The queue is not only commercial, it shows up directly in response time.
For a team weighing a production migration onto K3, the calculation changes shape. A strong model under rationing becomes a bet on the vendor’s capacity rather than a bet on the model itself.
Procurement teams tend to read that signal harshly. A vendor that cannot take an order today is hard to write into a contract that assumes growing volume over the next eighteen months.
The gradual reopening does soften the picture somewhat. Releasing slots in batches keeps quality of service stable for paying users, which is the sane call even though it caps growth in the short run.
Other Chinese labs attacked the same wall through optimisation rather than rationing, as when DeepSeek pulled 85 percent more speed out of its GPUs with DSpark. Moonshot does not have that card in hand today.
Splitting the offer in two also reveals the strategy. Isolating programming workflows in a dedicated plan lets the company steer its heaviest load separately, and reopen one door without reopening the other.
The signal reaching buyers stays mixed. Closing sales two days after launch proves the product has traction and undercuts the service guarantee any serious deployment requires.
More articles on Horizon
- Alibaba Ships Qwen 3.8, Its Biggest AI Model Yet
- Gemini 3.5 Pro Delayed Again as Google Rebuilds It
- Kimi K3 Closes the Gap With Top US AI Models
Open weights do not shrink the compute bill
The episode cuts against a common assumption. An open model is supposed to ease the compute burden on the lab that publishes it, since anyone can host it elsewhere.
In practice demand piles onto the service the vendor hosts. A 2.8 trillion parameter model needs infrastructure that very few organisations can stand up, and an openness pledge does nothing about that hardware reality.
The bottleneck stays in the hardware. Nvidia pushed its next rack back by a full year, a slip we unpacked when the next generation timeline slid out to 2028.
Rivals can read this two ways. Moonshot just proved the market appetite for a top-tier Chinese model, and exposed its own execution ceiling in the same breath.
The timing lands badly for Moonshot. Alibaba unveiled Qwen 3.8 the next day with a preview billed at 10 percent of the standard price, precisely as curious developers hit a closed door at its rival.
The deeper question runs past this one case. A frontier model is now judged as much on the industrial capacity of its vendor as on its scores, and that dimension appears on none of the published leaderboards.
The window favours whoever can serve without rationing. Every day the door stays shut sends the teams who planned to trial K3 this week toward a vendor able to absorb the load.
Follow the story on Horizon.


