DeepSeek has released V4.1-Flash, and the headline is unusual: the budget model beats the company’s own flagship.
Asymmetric Architecture
V4.1-Flash is the smallest model in DeepSeek’s new architecture family — a 552B-parameter MoE using a new Causal Encoder-Decoder design with just 8B active parameters for input and 16B for output. Combined with new pre-training methods and larger-scale RL post-training, DeepSeek says benchmark results land ahead of flagship models, including V4-Pro.
The efficiency story is dramatic: compared with the previous generation, the KV cache needs just 1/4 the HBM and 1/8 the SSD storage. Since cache-hit charges often dominate agent costs, compressing the cache cuts those bills significantly — a direct play for the agentic workloads that now drive API usage.
V4-Pro Is Being Phased Out
Third-party tests put V4.1-Flash ahead of V4-Pro on performance, cost, speed, and total runtime — so DeepSeek is retiring V4-Pro. From 04:00 UTC on September 14, 2026, all deepseek-v4-pro requests route to V4.1-Flash at V4.1-Flash rates, until a future V4.1-Pro launch. The model is live on the API now under the name deepseek-flash, with native multimodal support.
New pricing also took effect September 10: off-peak rates are 50% of peak, letting flexible workloads schedule around demand. Official partners WorkBuddy (including CodeBuddy) and OpenCode already support the model.
Open Source Commitment
Weights are available on Hugging Face along with the full technical report, and DeepSeek says it will work closely with the open-source community on inference support and additional deployment options — including large-scale deployments of 2,000+ GPU clusters.