Full Breakdown
Chinese AI Firms Grapple with Nvidia Chip Shortage Amid Soaring Inference Demand
8/21/2026, 8:18:13 AM
Inference Demand Outpaces Domestic Chip Capacity
Chinese artificial-intelligence companies are encountering acute compute constraints as the industry moves from model training to large-scale inference. While training still depends on high-end processors, many inference workloads can be adapted to domestic hardware. However, industry insiders say that complex, high-value tasks—most notably code generation—continue to require Nvidia GPUs. The shortage of these chips forces firms to rely on domestic processors that can only support lower-quality inference tiers, limiting commercial viability for premium services.
Context: Token Usage Surge and Shift to Inference
The National Data Administration reports that average daily token calls in China exceeded 140 trillion in March, a rise of more than 1,000-fold since early 2024. This explosion reflects AI models becoming more agentic, performing real-world tasks rather than merely answering queries. The surge in high-quality token demand intensifies pressure on the limited supply of Nvidia chips, creating a bottleneck for developers seeking to monetize advanced capabilities such as coding assistance.
Industry Response and Viability Concerns
Companies are therefore optimising software to stretch the scarce Nvidia supply while continuing to develop domestic alternatives for less demanding workloads.
Verbatim Quotes
- “The demand side is now showing a bipolarisation,” — Guan Jiawei, vice-president of inference optimisation start-up Approaching
- “If we rely solely on domestic chips for inference, they can only handle the low-quality tier – the tier with weak demand and weak monetisation,” — Guan
