Two stories about Chinese AI are true at the same time, and most coverage only tells one of them. The first: Chinese open-weight labs now dominate a major usage marketplace so completely that OpenAI and Google briefly vanished from its top 10 entirely. The second: the hardware underneath China's own frontier accelerator, Huawei's Ascend line, is capped by a memory bottleneck that no amount of model-design talent fixes. Both are backed by primary, dated numbers. Here's both stories, together, the way they actually interact.
The model landscape, July 2026
The open-weight tier split three ways this month rather than being led by one lab. Kimi K3 (Moonshot, 2.8 trillion parameters, MoE) leads the Arena.ai Frontend Code Arena, ahead of Claude Fable 5. GLM-5.2 (Zhipu, 744 billion parameters) tops the Artificial Analysis Intelligence Index among open weights and beats GPT-5.5 on several long-horizon coding benchmarks at roughly a sixth of the price. DeepSeek V4 Pro (1.6 trillion parameters, MIT-licensed) leads raw SWE-bench Verified among downloadable weights at 80.6%. MiniMax M3 (roughly 428 billion parameters) scores 80.5% on the same benchmark, a hair behind. Qwen3-Coder-Next is the strongest model that still runs locally on 46GB of memory.
| Lab / Model | Parameters | Notable claim (July 2026) |
|---|---|---|
| Moonshot / Kimi K3 | 2.8T (MoE) | Leads Arena.ai Frontend Code Arena, ahead of Claude Fable 5 |
| Zhipu / GLM-5.2 | 744B | Tops the open-weights Intelligence Index; beats GPT-5.5 on long-horizon coding at ~1/6 the price |
| DeepSeek / V4 Pro | 1.6T | Leads raw SWE-bench Verified (80.6%) among downloadable weights; MIT-licensed |
| MiniMax / M3 | ~428B | 80.5% SWE-bench Verified, ~1M-token context |
| Alibaba / Qwen3-Coder-Next | — | Best-performing model runnable locally on 46GB |
What the usage data actually shows
The clearest evidence that this isn't just benchmark bragging is OpenRouter's own traffic data. As of late June 2026, Chinese models accounted for 48% of all token processing on the platform against 20% for US models. By mid-July, Chinese models reached a record 58% of tokens specifically among those processed by US-headquartered firms on the platform, briefly spiking to 63% in the first week of July. DeepSeek alone is the single largest individual provider on OpenRouter by token volume, at 16.3% — and Chinese providers collectively (DeepSeek, Tencent, Xiaomi, MiniMax, Qwen) account for roughly 44% of top-10 token volume. In July, OpenAI and Google models were absent from OpenRouter's top 10 entirely; Anthropic was the sole US survivor. Companies reported switching for one clear, unglamorous reason: Chinese open models run 60-90% cheaper than flagship US alternatives, and per one industry estimate, open models now handle nearly a third of all tokens on the platform while accounting for under 4% of spend.
Worth being precise about which number means what here: "48% of all traffic," "58% of US-firm traffic," and "63% intraday" are three different, real, non-contradictory figures answering three different questions — treating any one of them as "Chinese models' overall global AI market share" overstates what OpenRouter, a single marketplace, actually represents.
The ceiling underneath the model race
Here's the part usage-share coverage tends to skip: model quality and usage share run on a completely different constraint than compute-at-scale. Huawei's Ascend 910C — the accelerator China would need in volume to reduce its own dependence on smuggled or diverted Nvidia hardware — is built on SMIC's 7nm process, roughly equivalent to a TSMC 2018-era node. For years, Huawei supplemented that domestic logic with a stockpile of foreign-made dies: reporting puts the total at roughly 2.9 million Ascend-class dies obtained from TSMC via shell companies before the October 2024 cutoff, a "die bank" that carried Huawei's production through 2024 and into 2025. As of early 2026, that stockpile is effectively exhausted — future Ascend production now depends entirely on SMIC's own wafers and domestic HBM packaging.
That's where CXMT comes in, and where the real ceiling sits. Per SemiAnalysis's Huawei Ascend production-ramp analysis, CXMT is expected to produce only around 2 million HBM stacks in 2026 — enough to support roughly 250,000 to 300,000 Ascend 910C-class accelerators for the entire year. That's the number that actually caps China's AI compute buildout at the hardware layer, regardless of how many parameters DeepSeek or Kimi ship, or how cheap their tokens get on OpenRouter.
The enforcement gap, and why it doesn't undo the ceiling
The export-control regime meant to prevent Chinese firms from working around this ceiling via smuggled Nvidia hardware has real, documented leaks. During a roughly year-long regulatory gap (May 2025 to May 2026), Chinese buyers reportedly used Singapore, Malaysia, and UAE holding companies to purchase Blackwell and H100-class GPUs without the licenses that would otherwise be required. One Singapore cloud provider imported $4.6 billion in Nvidia hardware — at least 136,000 GPUs by import records, against only 86,000 Nvidia itself catalogued as delivered there, leaving tens of thousands of chips unaccounted for. A June 5, 2026 seizure at Kuala Lumpur International Airport intercepted 72 servers loaded with advanced AI chips. Commerce and DOJ together announced nearly $420 million in penalties and forfeitures tied to semiconductor diversion over the trailing 12 months.
The BIS's May 31, 2026 guidance directly targets that gap: license requirements now extend to any entity whose ultimate parent is headquartered in China, Russia, or another Country Group D:5 nation, regardless of where that entity is physically incorporated — closing the specific Singapore/Malaysia shell-subsidiary channel used during the gap year. What it doesn't do is claw back the roughly 2.9 million dies Huawei already banked before October 2024, or add HBM capacity CXMT doesn't yet have. Diverted Nvidia chips supplement China's compute base at the margins; they don't change the arithmetic on Huawei's own accelerator line, which stays capped at a few hundred thousand units a year until domestic memory catches up — an 18-36 month problem at minimum, not something one enforcement action resolves.
Sources: SemiAnalysis: Huawei Ascend Production Ramp, TechPowerUp on the Ascend 910C's TSMC dies, Benzinga on OpenRouter's 58% record share, Yahoo Finance on the 46% enterprise figure, Datagravity (Chris Zeoli) on China's open-weight takeover, Morph on the July 2026 open-weight coding leaderboard, Tech Times on the BIS D:5 guidance, Tech Times on the export-control gap and diversion evidence.
