Ox Alpha showed up on OpenRouter on August 20 carrying a million-token context window, a bill of zero, and no company name attached to it. Three days on, half the industry is guessing where it came from while the other half has already wired it into a coding pipeline.
Key Takeaways
- Ox Alpha is a reasoning model built for coding and sustained agentic work, served free on OpenRouter by a provider that will not name itself.
- The spec sheet lists 1,048,576 tokens of context, 131,072 tokens of output, text, image and video input, and prompt retention on the provider side.
- Theories on its origin split between a Chinese lab close to Z.ai’s GLM models and Microsoft’s MAI programme, with hard evidence for neither.
Have an AI Sum Up This Article
ChatGPTA reasoning model shipped with nobody’s name on it
The listing describes a reasoning model built for coding, long-running agentic work and production workloads. Whoever built it sits behind the stealth label, a third party that chooses not to identify itself for the duration of the preview.
None of that is unprecedented. Masked models pass through the platform fairly regularly ahead of an official launch, giving whoever built them a window to gather real usage feedback at scale without putting a brand behind a result they cannot yet predict.
What is unusual here is how much traction it picked up in three days. Stripe chief executive Patrick Collison tried it and called it very impressive, and Reddit threads and X timelines have filled with amateur evaluations hunting for a tell about who is behind it.
The pull comes from a rare combination. A coding and agents positioning, a million-token window and a zero invoice, arriving exactly when the fastest coding agents on the market are also the most expensive ones to run.
Ox Alpha therefore lands as a loss leader in a segment where invoices climb fast. For a provider preparing a launch, giving inference away for a few weeks is a customer acquisition line, not a pricing mistake.
The free window carries an implicit expiry date all the same. A model this size served at 24 tokens per second against a stated 99.99% uptime burns GPU continuously, and nobody funds that indefinitely out of goodwill.
A million tokens of context, and your prompts kept
The technical listing published on the platform puts up top-tier numbers: 1,048,576 tokens of context, 131,072 tokens of maximum output, text, image and video accepted on the way in, text only on the way out.
Function calling works through the standard tools and tool_choice parameters, and the model returns JSON on request without schema validation. That is enough to drop it into an existing agent chain without rewriting the tooling layer around it.
No token is billed, inbound or outbound. For a team running agents across several hours at a stretch, the offer is blunt: the single largest cost line disappears for the length of the preview.
Throughput is the number worth reading twice before planning around it. At 24 tokens per second, a long agent run that fills a large share of that context window takes real wall-clock time, which suits overnight batches better than an interactive loop a developer sits and waits on.
The trade sits in plain sight. Prompts and completions are retained by the provider, outside of training use under the stealth model terms, but retained nonetheless, by a party whose name and jurisdiction nobody knows.
For prototyping, for internal benchmarks, or for work on a repository that is already public, that trade is easy enough to justify. For proprietary code or customer data, routing requests to an unknown operator that archives them is a professional failure, whatever the model scores.
That is the line every team has to draw this week, and benchmark tables do not help with it. What matters is what leaves your servers. The rest of the market is moving the same way with named guarantees attached, since DeepSeek just pushed its V4-Flash-Vision close to Opus 4.8 while capping every image at 384 tokens.
More articles on Horizon
- DeepSeek V4-Flash-Vision Closes In on Opus 4.8
- Claude Security Now Scans Code With Mythos 5
- OpenAI Closes the Gap With Anthropic in Business
Z.ai or Microsoft, the two theories splitting the testers
Which leaves the question driving every thread: who pays for the GPUs. Some testers say the output style reads like Z.ai’s GLM family, which would place Ox Alpha in the Chinese lab lineage. Others point at Microsoft’s MAI programme instead.
Neither camp has produced anything decisive, and the two theories tell opposite stories. A Chinese lab would extend the run of Asian models undercutting American pricing, the way MiniMax M3 cleared GPT-5.5 on open weights earlier in the year.
A masked Microsoft would say something else entirely. It would show an American giant testing whether it can stand up a frontier model outside the shadow of OpenAI, letting the market judge the work without the lift or the drag of its own brand.
On the competitive side the effect is immediate and does not wait for the answer. A free preview aimed at coding and agents squeezes every paid offer in the same bracket, including the ones that recently cut prices to stay in the comparison.
Every extra week of anonymity also extends a live trial the market is populating for free. Whenever the name lands, real-world performance will already be documented by thousands of testers, flattering or not, without a line of launch budget spent.
The gamble cuts both ways. A preview that disappoints buries the launch before the announcement, and the anonymity protecting the provider today will block any attempt to change minds that have already settled.
Our read comes down to two lines. The free window is worth using to evaluate Ox Alpha on workloads that can leak without damage, and the code that matters stays away from a faceless server until that face exists.
Follow the story on Horizon.


