TL;DR
DeepSeek’s V4-Flash became the world’s most-used AI model by token volume, hitting 7.1 trillion tokens weekly. Priced at $0.14/M input and $0.28/M output tokens — over 100x cheaper to run than Anthropic’s Claude Fable 5 — the model established a new industry cost benchmark. Notably, nine of the top ten models by global token usage are now Chinese.
The Numbers
DeepSeek’s V4-Flash achievements:
- Global token usage: 7.1 trillion tokens per week — #1 worldwide
- Pricing: $0.14/M input, $0.28/M output tokens
- Cost advantage: Over 100x cheaper to run than Claude Fable 5
- Agent benchmarks: 82.7% on Terminal-Bench, beating its own 1.6T Pro model
- Chinese dominance: 9 of top 10 models by token usage are now Chinese
The model officially exited preview on July 31, 2026 as “V4-Flash 0731” under an MIT license — free to self-host or access via API at rock-bottom prices.
Why Chinese Models Are Winning Usage
Several factors explain the Chinese dominance:
- Price: Chinese labs have driven API costs down to fractions of US pricing
- Open weights: MIT and Apache licenses allow self-hosting without per-token costs
- Capability parity: Chinese models now match US models on many benchmarks
- Developer economics: For high-volume workloads, cost dominates model choice
The cost difference is stark: Claude Fable 5’s pricing makes it a premium product, while V4-Flash’s pricing makes it the default for volume workloads.
Industry Implications
- US pricing pressure: OpenAI’s 80% Luna price cut was a direct response to this trend
- Commoditization: AI inference is becoming a commodity market where price dominates
- Revenue pressure: US labs face margin compression as Chinese competitors undercut
- Enterprise adoption: Cheap models accelerate AI adoption but reduce vendor revenue
The US Response
OpenAI cut GPT-5.6 Luna prices by 80% to $0.20/M input on July 30. Google reduced Gemini 3.6 Flash output costs by 17%. Anthropic’s Sonnet 5 introduced pricing at $2/$10 per million.
But the gap remains: V4-Flash at $0.14/M input is still cheaper than any Western flagship, and its open-weight MIT license means it can be run for zero per-token cost on self-hosted infrastructure.
For the AI industry, the token usage rankings represent a fundamental shift: the center of gravity in AI usage is moving toward cheap, open Chinese models — and price wars are reshaping the economics of the entire industry.