DeepSeek V4.1-Flash: Smarter, Faster, Cheaper — and It Beats Their Own Flagship

DeepSeek releases V4.1-Flash, a 552B asymmetric MoE that outperforms V4-Pro on benchmarks at lower cost, with a KV cache needing just 1/4 the HBM of the previous generation.

Thursday September 10, 2026 Source: DeepSeek
TL;DR — Quick Answer

DeepSeek released V4.1-Flash, a 552B-parameter asymmetric MoE with just 8B active input and 16B output parameters, natively multimodal. Its KV cache needs only a quarter of the HBM of the previous generation. Third-party tests put V4.1-Flash ahead of the flagship V4-Pro on performance, cost, speed, and total runtime — DeepSeek is phasing V4-Pro out, and all deepseek-v4-pro requests will route to V4.1-Flash at Flash pricing from September 14, 2026.

Key Takeaways

DeepSeek V4.1-Flash: Smarter, Faster, Cheaper — and It Beats Their Own Flagship — AI news article illustration

DeepSeek has released V4.1-Flash, and the headline is unusual: the budget model beats the company’s own flagship.

Asymmetric Architecture

V4.1-Flash is the smallest model in DeepSeek’s new architecture family — a 552B-parameter MoE using a new Causal Encoder-Decoder design with just 8B active parameters for input and 16B for output. Combined with new pre-training methods and larger-scale RL post-training, DeepSeek says benchmark results land ahead of flagship models, including V4-Pro.

The efficiency story is dramatic: compared with the previous generation, the KV cache needs just 1/4 the HBM and 1/8 the SSD storage. Since cache-hit charges often dominate agent costs, compressing the cache cuts those bills significantly — a direct play for the agentic workloads that now drive API usage.

V4-Pro Is Being Phased Out

Third-party tests put V4.1-Flash ahead of V4-Pro on performance, cost, speed, and total runtime — so DeepSeek is retiring V4-Pro. From 04:00 UTC on September 14, 2026, all deepseek-v4-pro requests route to V4.1-Flash at V4.1-Flash rates, until a future V4.1-Pro launch. The model is live on the API now under the name deepseek-flash, with native multimodal support.

New pricing also took effect September 10: off-peak rates are 50% of peak, letting flexible workloads schedule around demand. Official partners WorkBuddy (including CodeBuddy) and OpenCode already support the model.

Open Source Commitment

Weights are available on Hugging Face along with the full technical report, and DeepSeek says it will work closely with the open-source community on inference support and additional deployment options — including large-scale deployments of 2,000+ GPU clusters.


Frequently Asked Questions

What is DeepSeek V4.1-Flash?

V4.1-Flash is DeepSeek's newest model, released September 10, 2026. It is a 552B-parameter mixture-of-experts model with a new asymmetric Causal Encoder-Decoder architecture — 8B active parameters for input and 16B for output — and native visual understanding.

Does V4.1-Flash beat V4-Pro?

Yes. Tests by multiple parties put V4.1-Flash ahead of DeepSeek's flagship V4-Pro on performance, cost, speed, and total runtime. DeepSeek is phasing out V4-Pro.

What happens to deepseek-v4-pro API requests?

Starting at 04:00 UTC on September 14, 2026, all deepseek-v4-pro requests route to V4.1-Flash and are billed at V4.1-Flash rates. This continues until V4.1-Pro launches. V4-Flash and V4-Flash-Vision-Exp are retired, with legacy names temporarily routing to V4.1-Flash.

How efficient is V4.1-Flash's memory usage?

Compared with the previous generation, V4.1-Flash's KV cache needs just 1/4 the HBM and 1/8 the SSD storage — significantly cutting agent costs where cache-hit charges dominate.

Is V4.1-Flash open weight?

Yes. Model weights are on Hugging Face at huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash along with the technical report PDF, and DeepSeek says it is working with the open-source community on inference support.

This article is based on the official announcement from DeepSeek . Read the original for full technical details.

Related Articles

Back to all news