← pub

AI Agents Now Outconsume Humans in Tokens, Fueling a Shift Toward Cheaper Chinese Models

Autonomous AI agents now generate five times more tokens than humans, pushing costs toward cheap Chinese models even as US labs still command most AI revenue.

On the surface, the numbers read like a coronation: agentic-token volume on the marketplace OpenRouter hit 7.3 trillion tokens in a single day this August, fourteen times the level six months earlier, and Chinese developer DeepSeek’s V4-Flash undercuts Anthropic’s Claude Fable 5.1 on price by roughly 70 to 1.

Revenue hasn’t followed.

Leading Chinese AI firms reported 2025 revenue of between 100 million yuan (US$14.6 million) and 700 million yuan, against more than $25 billion for OpenAI and $30 billion for Anthropic.

Call it triage. Confronted with a caseload agents generate faster than any human queue could process, companies are increasingly sorting workloads by urgency rather than sending everything to the most expensive specialist: routine, repetitive tasks go to lower-cost models, reserving frontier systems for complex reasoning, an approach Sun Wenhao, a senior investment manager at TH Capital, said keeps AI bills manageable as usage keeps climbing.

The logic has a precedent: nationwide enthusiasm for the open-source agent OpenClaw pushed token consumption up roughly tenfold in about ten weeks, Infinigence co-founder Xia Lixue said, straining a system with no ward built for that volume.

Wall Street reads the triage differently: not as care, but as competition for market share. Morgan Stanley analysts outlined three ways the contest between open-weight and closed models could resolve, from an open-weight sweep that pushes prices down further to a closed-model oligopoly that keeps them elevated, pointing to the “Jevons Paradox” as evidence that cheaper access tends to enlarge total demand rather than shrink the market.

A separate MIT Sloan School of Management study, though, found agents can produce code, analysis and documents faster than employees can vet them, suggesting the real chokepoint on how far this triage can run isn’t processing speed. It’s who’s left to check the chart.

The Volume Trap That’s Everyone’s Problem

Zoom in on the numbers above and the gap looks like arithmetic: which price per million tokens wins. Zoom out to a single desk and it’s a workflow problem, one specific enough to already have a name — cheap generation inflates a task’s periphery while what counts as “done” never moves. I personally experienced the token boom and some of the mess it leaves behind.

Zoom out to an organization and it’s a budgeting problem: one company reportedly burned through roughly $500 million in Claude credits in a single month after skipping usage guardrails.

Zoom out once more and the frame changes kind, not size: whether an organization can govern work it delegates to a tool that, when its own instruction-state breaks, can’t explain the failure even when asked directly.

In my earlier piece, I talked about separate symptoms:

  • Uunchecked spend
  • Retrieval failures that misattribute or truncate a source
  • A memory bug that lets an agent’s own global instructions silently revert without anyone, including the agent, noticing

Each sits in a different domain (in this case, cost, retrieval, session state) but the shape underneath doesn’t change. It bears repeating that tokens spent was never a proxy for value produced, any more than lines of code were ever a proxy for working software.

The news above is this same recursive pattern with an international-market-sized blast radius. Cheap tokens pull usage toward Chinese models on the cost side; cheap tokens pull usage past the point of review on the governance side. Two symptoms projecting the same silhouette.


Aklatan’s news and analysis drills down to the structural mechanics, geopolitical shifts, and hidden constraints truly driving AI and Asian tech ecosystems and knowledge work.

See coverage span here: Aklatan’s News and Analysis

Generative AI Transparency:

This article was written primarily with generative AI, specifically SupraGraphos’ A.C.E. News Module. Reviewed with human post-editing, all sources and claims are confirmed as of the time of writing.