China's daily AI token calls reached 140 trillion by March 2026, a more than 1,000-fold increase from roughly 100 billion in early 2024, according to figures cited by Wei Liang, deputy head of the China Academy of Information and Communications Technology (CAICT), in a CCTV Finance program preview published on July 182.
The surge reflects the scale at which Chinese companies are now consuming inference compute. A Beijing-based ByteDance employee told the South China Morning Post that he burns through close to a billion AI tokens every month — and that volume "doesn't even put me near the top of the consumption rankings in my department"1. "Using AI has become part of almost every task now," the employee said, speaking on condition of anonymity because he was not authorized to comment publicly.
The CAICT-cited report linked the increase to wider adoption of AI agents, in which a single user instruction can trigger multiple model calls, compounding token demand well beyond what direct human prompting would generate. That agentic multiplier effect has in turn created demand for token pricing and scheduling systems to manage the load.
ANALYSIS The 1,000-fold figure in roughly two years underscores how quickly inference — rather than training — is becoming the dominant compute workload in China's AI ecosystem. The ByteDance anecdote illustrates the mechanism: when AI is embedded in routine workflows across large engineering organizations, per-employee token consumption can reach volumes that would have been enterprise-scale totals a short time ago.
The emergence of token pricing and scheduling infrastructure points to a maturing supply chain around inference delivery, where raw model availability is no longer the bottleneck — cost management and orchestration are.