DeepSeek V4-Flash Tops Global Token Usage: 7.1 Trillion Tokens Weekly, 9 of Top 10 Models Now Chinese

DeepSeek's V4-Flash at $0.14/M input and $0.28/M output tokens topped global token usage at 7.1 trillion tokens weekly. A research firm found it over 100x cheaper to run than Claude Fable 5, with nine of the top ten models now Chinese.

Tuesday August 4, 2026 Source: news.cgtn.com
TL;DR — Quick Answer

DeepSeek''s V4-Flash became the world''s most-used AI model by token volume, hitting 7.1 trillion tokens weekly. Priced at $0.14/M input and $0.28/M output tokens — over 100x cheaper to run than Anthropic''s Claude Fable 5 — the model established a new industry cost benchmark. Notably, nine of the top ten models by global token usage are now Chinese.

Key Takeaways

DeepSeek V4-Flash Tops Global Token Usage: 7.1 Trillion Tokens Weekly, 9 of Top 10 Models Now Chinese — AI news article illustration

The Numbers

DeepSeek’s V4-Flash achievements:

The model officially exited preview on July 31, 2026 as “V4-Flash 0731” under an MIT license — free to self-host or access via API at rock-bottom prices.

Why Chinese Models Are Winning Usage

Several factors explain the Chinese dominance:

The cost difference is stark: Claude Fable 5’s pricing makes it a premium product, while V4-Flash’s pricing makes it the default for volume workloads.

Industry Implications

The US Response

OpenAI cut GPT-5.6 Luna prices by 80% to $0.20/M input on July 30. Google reduced Gemini 3.6 Flash output costs by 17%. Anthropic’s Sonnet 5 introduced pricing at $2/$10 per million.

But the gap remains: V4-Flash at $0.14/M input is still cheaper than any Western flagship, and its open-weight MIT license means it can be run for zero per-token cost on self-hosted infrastructure.

For the AI industry, the token usage rankings represent a fundamental shift: the center of gravity in AI usage is moving toward cheap, open Chinese models — and price wars are reshaping the economics of the entire industry.

Frequently Asked Questions

What model tops global token usage?

DeepSeek''s V4-Flash leads with 7.1 trillion tokens processed per week.

How much cheaper is V4-Flash than Claude Fable 5?

Over 100x cheaper to run, at $0.14/M input and $0.28/M output tokens.

Why are Chinese models winning on usage?

Cheaper pricing, open-weights licenses like MIT and Apache, capability parity, and favorable developer economics.

This article is based on the official announcement from news.cgtn.com . Read the original for full technical details.

Related Articles

Back to all news