AI Research #DeepSeek#V4-Flash#China#token usage#pricing#AI models#price war#global rankings

DeepSeek V4-Flash Tops Global Token Usage: 7.1 Trillion Tokens Weekly, 9 of Top 10 Models Now Chinese

DeepSeek's V4-Flash at $0.14/M input and $0.28/M output tokens topped global token usage at 7.1 trillion tokens weekly. A research firm found it over 100x cheaper to run than Claude Fable 5, with nine of the top ten models now Chinese.

Tuesday August 4, 2026
DeepSeek V4-Flash Tops Global Token Usage: 7.1 Trillion Tokens Weekly, 9 of Top 10 Models Now Chinese

TL;DR

DeepSeek’s V4-Flash became the world’s most-used AI model by token volume, hitting 7.1 trillion tokens weekly. Priced at $0.14/M input and $0.28/M output tokens — over 100x cheaper to run than Anthropic’s Claude Fable 5 — the model established a new industry cost benchmark. Notably, nine of the top ten models by global token usage are now Chinese.

The Numbers

DeepSeek’s V4-Flash achievements:

The model officially exited preview on July 31, 2026 as “V4-Flash 0731” under an MIT license — free to self-host or access via API at rock-bottom prices.

Why Chinese Models Are Winning Usage

Several factors explain the Chinese dominance:

The cost difference is stark: Claude Fable 5’s pricing makes it a premium product, while V4-Flash’s pricing makes it the default for volume workloads.

Industry Implications

The US Response

OpenAI cut GPT-5.6 Luna prices by 80% to $0.20/M input on July 30. Google reduced Gemini 3.6 Flash output costs by 17%. Anthropic’s Sonnet 5 introduced pricing at $2/$10 per million.

But the gap remains: V4-Flash at $0.14/M input is still cheaper than any Western flagship, and its open-weight MIT license means it can be run for zero per-token cost on self-hosted infrastructure.

For the AI industry, the token usage rankings represent a fundamental shift: the center of gravity in AI usage is moving toward cheap, open Chinese models — and price wars are reshaping the economics of the entire industry.

Back to all news