Agent tokens on OpenRouter went from 0.51 trillion to 7.3 trillion since February, while human consumption grew only 2.8x over the same window. The platform puts the crossover on February 6, the day machines most likely overtook the people using them.
Key Takeaways
- Agentic consumption has multiplied fourteen times since February 6, 2026.
- Humans grew 2.8x over the same window, roughly five times slower.
- About 70% of agentic tokens come from cached prompts, billed at reduced rates.
Have an AI Sum Up This Article
ChatGPTFebruary 6 Passed Without Anyone Noticing
Public readings from OpenRouter show a curve that splits in two early this year. Agent tokens climbed from 0.51 trillion to 7.3 trillion since February 6, a fourteenfold jump across roughly six and a half months. Nothing in the underlying model pricing moved enough to explain a curve of that shape.
Over exactly the same window, human consumption grew 2.8x. The gap in slope is the real signal rather than the absolute level, because it says the two kinds of usage no longer answer to the same growth driver. Human demand scales with headcount, and machine demand scales with whatever a team decides to automate next.
Peter Walker, an analyst at OpenRouter, put the reading in one line alongside the platform’s public consumption rankings: February 6, 2026 may have been the last day humans consumed more tokens than AI agents did.
The mechanism behind the split sits in how an agentic run is shaped. A human query returns one answer, while a task handed to an agent fires a chain of model calls, tool uses and reasoning steps before anything comes back. A single instruction can therefore cost what a dozen conversations used to.
One caveat belongs on the scope. These readings cover traffic routed through a single aggregation platform rather than the whole inference market, and nothing guarantees that calls made straight to lab endpoints follow the same trajectory.
A multiplier effect sits on top. Autonomous systems spawn further processes of their own, and each one consumes in turn, so growth now tracks the number of delegated tasks rather than the number of people logged in.
Why the Bill Is Not Climbing at the Same Rate
The raw number gives a misleading impression of cost. Close to 70% of agent tokens come from prompts served out of cache, which most providers bill at a sharply reduced rate.
The reason is mechanical. An agent repeating a task resends the same system context, the same instructions and often the same reference documents on every loop, which is exactly the profile caching was designed to absorb. The more repetitive the automation, the cheaper each additional loop becomes.
For an engineering team, the practical consequence is that tracking raw token volume stops working as an alarm. Two workloads identical in volume can produce wildly different invoices depending on reuse, a trap finance teams tend to discover at month end.
Optimisation therefore changes shape. Cutting the bill is no longer about writing shorter prompts, it is about keeping context stable enough to stay cacheable from one call to the next, a discipline token saving guides already spell out.
The gap between those two readings feeds a public misunderstanding. The raw curve travels because it looks dramatic, while the only figure a technical director can act on is cost per completed task, and no public dashboard publishes that today.
Still, 70% cached also means 30% billed at full rate on a volume multiplied by fourteen. Absolute spend goes up, just more slowly than the headline curve suggests.
More articles on Horizon
- Wan 3.0 Builds Thirty Seconds of Video From a PDF
- V4-Flash-Vision Test: DeepSeek Wins on Agent Work
- Ox Alpha: A Free AI Model Arrives With No Vendor
Inference Providers Are Changing Customers
If the trend holds, the typical customer of an inference provider is no longer a person in front of an interface but a program running around the clock. The two profiles share neither peak patterns, latency requirements nor price sensitivity.
Human usage follows an office rhythm, with quiet nights and slow weekends. A fleet of agents flattens the load across twenty four hours, which lifts utilisation on compute estates and makes capacity planning far more predictable. Idle capacity at three in the morning stops being dead weight.
On the supply side, the expected move lands on pricing sheets. A provider whose volume arrives mostly as repeated prompts has every reason to push harder cache discounts and capture agentic workloads before its rivals do.
Labs are exposed unevenly depending on their mix. A model that dominates consumer chat captures a 2.8x growth curve, while a model well positioned on agentic coding captures the fourteenfold one.
The aggregator position gains from all of this. Whoever watches both kinds of traffic pass through holds a view neither the labs nor their customers have, and that visibility is worth as much as the margin taken on each routed call.
Measurement itself becomes a question. Agent tokens and human tokens no longer carry the same economic value per unit, and rankings that add them into a single column now hide two separate markets.
Follow the story on Horizon.


