The Coinbase Chinese AI turn was confirmed by Brian Armstrong. Engineering teams now burn more tokens than ever on GLM 5.2 and Kimi 2.7, while the overall AI bill has dropped to roughly half of what Western models used to cost.
Key Takeaways
- Brian Armstrong confirms the switch to GLM 5.2 and Kimi 2.7, two Chinese models, for the bulk of Coinbase AI workloads.
- 91 percent of Coinbase developers never hit their previous usage caps, proving the migration did not constrain anyone.
- Cache hit rate jumped from 5 to 60 percent, and automatic routing by task, price and reuse handles the rest of the drop.
Have an AI Sum Up This Article
ChatGPTBrian Armstrong announces the switch
Coinbase CEO Brian Armstrong went public on social media to confirm the Coinbase Chinese AI move. The company now routes most of its AI workloads to GLM 5.2 and Kimi 2.7, both published by Chinese labs.
The headline number fits in a single sentence. Coinbase consumes more tokens than ever before, and pays around half of what it used to with Western providers. Developer usage keeps climbing while the budget curve stays flat.
The friction argument is just as loud. 91 percent of Coinbase developers never hit their previous monthly limits, which means the switch cost zero visible productivity. Nobody walked into Armstrong’s office to demand a return to US-built models because their workflow felt throttled.
The Coinbase context shapes the call. The group has just spent months publishing structural positions on infrastructure, including the quantum exposure of seven million bitcoins, and tight cost discipline on backend infrastructure has become a reflex.
Armstrong’s announcement is not a political statement. It is an operator decision, taken after benchmarks, presented as a budget read-out. The signal it sends to the rest of the market is nonetheless heavy.
Routing by task, cache from 5 to 60 percent: the technical kitchen
The Coinbase switch is not just a vendor swap. The team rolled out an automatic router that picks the model based on task, per-token price and caching potential.
The most striking gain sits on the cache layer. Hit rate moved from 5 to 60 percent, which means most calls no longer hit the remote model and instead serve a paid-for response from local storage. At constant volume, this single lever cuts costs by a factor few LLM optimizations ever reach.
Coinbase also chose to expose token consumption to every developer, without imposing hard caps. Each team sees what it spends, and the spend must map to a measurable business impact. The shift is cultural as much as technical, AI usage becomes a tracked budget line rather than an invisible commodity.
Routing by task reframes the debate. The market had settled into picking a premium model by default. Coinbase reminds everyone that a large share of production tasks runs perfectly on a mid-tier model, provided the selection is orchestrated upstream.
The operational result suggests that dependence on Western labs was partly a dependence by convenience. A well-built orchestrator suddenly widens the pool of acceptable providers.
Also on Horizon:
- Anthropic Survey: AI Does Half Your Work
- AI Political Bias: Only Gemini Stays Balanced
- Mythos Rollout: Trump Reopens Access to 100 US Partners
Chinese pricing puts OpenAI and Anthropic on edge
In the short term, the Coinbase Chinese AI move weighs on the contract renewals currently on the table at OpenAI and Anthropic. When a player this size publishes a 50 percent cut, every enterprise customer holds a fresh argument facing their account manager.
Other Western firms have already moved. Lindy switched to DeepSeek v4 per its CEO, and Snowflake is testing Chinese models as a direct alternative to OpenAI and Anthropic. The Coinbase proof point is likely to lengthen that list.
The presence of GLM 5.2 and Kimi 2.7 inside Western production stacks confirms a trend that has been building for months. MiniMax M3 had already opened the door by beating GPT-5.5 on several open benchmarks, without the rubber stamp of a US lab.
Over a three-to-six-month horizon, the price war between OpenAI and Anthropic turns mechanical. GPT-5.6-Sol has already been positioned by OpenAI as a token-efficiency response to Claude variants, and Chinese players now play the third party that forces the floor lower.
The political layer remains. Chinese models sit outside US jurisdiction, and any administration can decide overnight to shut the door. But as long as the price gap holds, the Coinbase Chinese AI calculation remains arithmetically defensible, and every Western CFO knows it.
Follow the story on Horizon.


