Moonshot has released the open weights of Kimi K3, its most capable model, which means anyone can now download all 2.8 trillion parameters for free. The package weighs 1.4TB in MXFP4 quantization, making it the largest open-weight release in history, and Kimi K3 becomes the first open model to lead a major long-horizon coding benchmark along the way.
Key Takeaways
- Kimi K3 weights are now free to download, 2.8 trillion parameters across 1.4TB of files.
- The model scores 42 on SWE Marathon, ahead of Fable 5 and GPT-5.6 Sol, a first for an open model.
- The API still bills 3 dollars per million input tokens and 15 for output, but free weights change the whole math.
Have an AI Sum Up This Article
ChatGPTThe biggest open model ever released fits in 1.4TB
The shift is clean. Moonshot first announced Kimi K3 without shipping the weights, a caution already visible when the lab halted Kimi sales as its GPUs maxed out. The weights are public now, and the model moves from a closed API to a building block anyone can host.
The numbers set the scale. Kimi K3 carries 2.8 trillion parameters, a one-million-token context window, native vision, and a thinking mode you can dial by effort. The full package weighs 1.4TB thanks to MXFP4 quantization, a format that compresses the weights without breaking quality, and the files sit on the lab’s official page over at Hugging Face.
This is not a first attempt. Kimi K3 already made noise when the open model closed the gap with top US systems, but the weights stayed out of reach. Today’s release shuts that access gap, and it lands in a Chinese summer already crowded with open models.
Context matters here. The Chinese camp is pushing hard on open, a move that picked up when Alibaba shipped Qwen 3.8, its biggest model yet, and Kimi K3 now sets the size bar above everything else.
42 on SWE Marathon, a first for an open model
The score that flips the story is SWE Marathon, a coding test built around long tasks. Kimi K3 lands 42 there, above Fable 5 at 35 and GPT-5.6 Sol at 39. For the first time, an open model leads a major long-horizon coding benchmark, a field closed models used to own outright.
The other readings confirm the level. The model scores 67.5 with its own KimiCode agent and 67.3 with the mini-SWE-agent harness, two measures that judge how well it carries a task end to end. Independent testers at Artificial Analysis rank it fourth of 189 models with an index of 57, at the very top of the open field.
For a team that builds, the point is not the podium but the independence. A model of this caliber that you host yourself removes the dependence on a third-party API, data control moves back to the client, and the marginal cost of a call drops to GPU electricity. It is a trade-off every player was already reopening when DeepSeek V4 retired its old model names to push prices down.
The nuance is hardware. Running 2.8 trillion parameters demands a serious GPU fleet, out of reach for a single workstation. So free weights mostly help hosting providers, specialized clouds, and large accounts that bring inference in house, not yet the lone developer on a laptop.
More articles on Horizon
- ChatGPT Voice Now Controls Your Whole Desktop
- Claude Opus 5 Matches the Best AI at Half the Cost
- Gemini Closes In on One Billion Monthly Users
A price floor that pressures OpenAI and Anthropic
A frontier-grade model that is free to host does not stay a Chinese story. It sets a price floor that closed models have to face, since a client can always weigh an API bill against the cost of a self-hosted deployment. Kimi K3’s API stays paid, at 3 dollars per million input tokens and 15 for output, with 30 cents for cached input.
The pressure hits the closed high end first. OpenAI and Anthropic sell capacity by the token, and an open rival that leads a coding benchmark narrows their room on price. Each lab now has to justify its rate with something other than raw quality alone, today reliability, ecosystem, or safety.
The open camp, meanwhile, is thickening fast. Kimi K3 arrives after a wave of free releases, right as Zhipu shipped China’s free rival to Claude, and these models together end up forming a credible alternative for anyone who wants off US APIs. The next signal to watch is how the closed labs move on price in the coming weeks.
Then there is data sovereignty. Hosting a Chinese model locally eases the fear of leaking to a foreign API, but shifts the debate to auditing the model itself. It is the kind of trade-off technical leaders will settle case by case, and Kimi K3 now hands them the means to do it.
Follow the story on Horizon.


