DeepSeek releases V4.1-Flash, a 552B MoE model with reduced API prices and strong benchmarks
DeepSeek has released DeepSeek-V4.1-Flash, a 552-billion-parameter Mixture-of-Experts model with a new Encoder-Decoder architecture and native visual understanding. The model activates 8B parameters for input and 16B for output, supports a 1 million token context window, and reduces KV cache to one-quarter of previous versions. API prices have been reduced, with off-peak pricing at $0.14 per million input tokens and $0.56 per million output tokens. The model outperforms competitors like GPT-5.6 Sol on several benchmarks.
Editorial responsibility
- No named human review is recorded for this page.
- Reports are grouped by semantic similarity and deterministic rules. Language models may assist titles, summaries, translation and cross-source analysis; the page reads the event directly, while its address stays stable when the title changes.
- Summary covers the current reports
Cross-source coverage
Reporting timeline
DeepSeek 4.1 Flash matches or beats GPT-5.6 Sol on benchmarks at 94% lower token price
DeepSeek has released its 4.1 Flash model, claiming it matches or outperforms OpenAI's GPT-5.6 Sol on several coding and agent benchmarks at a fraction of the token price. According to DeepSeek's tests, the new model achieved 74.2% on the DeepSWE software engineering benchmark, compared to GPT-5.6 Sol's 73.0% and the previous V4-Flash's 54.4%. DeepSeek also significantly reduced pricing versus V4-Flash: 32% for fresh input, 57% for cached input, and 9% for output. At peak rates, Flash charges $1.20 per million output tokens versus Sol's $20, a 94% reduction. Off-peak, the price drops to $0.60. The model features a new architecture that makes reading large inputs cheaper and uses sharing and compression of stored context to reduce memory requirements, making it particularly useful for agents repeatedly reading code, documents, and tool results. This move signals DeepSeek's aggressive return to the AI pricing war.
Read sourceDeepSeek releases V4.1-Flash with new Causal Encoder-Decoder architecture and lower API prices
DeepSeek has released V4.1-Flash, a new AI model featuring a Causal Encoder-Decoder architecture with native visual understanding, just six weeks after its July update to V4-Flash. The July release had improved post-training while keeping the architecture unchanged. V4.1-Flash has 552B MoE parameters, with 8B active during input processing and 16B during output generation. The company reports that KV-cache requirements are reduced to one-quarter of the HBM and one-eighth of the SSD storage compared to the previous generation, along with lower API prices. DeepSeek states that Flash now outperforms V4-Pro in capability, cost, and speed. Starting September 14, V4-Pro API requests will temporarily be routed to V4.1-Flash until V4.1-Pro becomes available. The post highlights that these significant leaps were achieved in just a few weeks through post-training, reflecting a new reality of weekly releases delivering substantial improvements.
Read sourceDeepSeek Releases V4.1-Flash with Causal Encoder-Decoder and Visual Understanding
DeepSeek has released V4.1-Flash, the smallest model in its new architecture family, featuring a novel Causal Encoder-Decoder architecture and native visual understanding capabilities. This 552-billion-parameter Mixture-of-Experts (MoE) model activates 8 billion parameters for input processing and 16 billion for output generation. Compared to the previous generation, KV cache requirements have been reduced to one-quarter of HBM and one-eighth of SSD storage, leading to lower API pricing. The article also outlines architectural differences from July's V4-Flash, details parameter and cache metrics, and explains temporary routing arrangements for the V4-Pro API, enabling a quick assessment of the update's substance.
Read sourceShow 4 older updatesHide older updates
DeepSeek AI Releases DeepSeek-V4.1-Flash With 1M Context and Reduced KV Cache
DeepSeek AI has released DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model featuring a 552B-parameter backbone and 196B Engram parameters. The model supports a 1 million token context window and activates 8B parameters during prefill and 16B during decode. A key innovation is the reduction of global KV cache to 890 bytes per token, which is approximately one-quarter of the cache used by DeepSeek-V4-Flash and one-four-hundred-thirty-seventh of DeepSeek-V1. The article details techniques for KV cache compression and pre-pruning architecture, enabling evaluation of long-context deployment costs. This release represents a significant advancement in efficient long-context AI model deployment.
Read sourceDeepSeek releases V4.1 Flash, a 552B-parameter MoE model on Hugging Face
DeepSeek has released its V4.1 Flash model on Hugging Face, a 552-billion-parameter Mixture-of-Experts (MoE) model featuring a new Encoder-Decoder architecture. The model employs a novel pre-training method and has undergone extensive reinforcement learning post-training. It is described as the smallest model in DeepSeek's new architecture family and includes native visual understanding capabilities. The announcement highlights the model's availability on the Hugging Face platform, making it accessible to the AI research and development community.
Read sourceDeepSeek V4.1 Flash released: 552B MoE model outperforms GPT-5.6 Sol on coding benchmarks
DeepSeek has released V4.1 Flash, a 552 billion parameter Mixture-of-Experts (MoE) model that activates only 8 billion parameters for input and 16 billion for output. The model is outperforming competitors such as GPT-5.6 Sol and Opus 5 on several agentic and coding benchmarks. Specific scores include 88.1 on CyberGym (vs. 84.5 for GPT-5.6 Sol), 54.8 on Automation-Bench (vs. 45.8), and 74.2 on DeepSWE (vs. 73.0). Off-peak API pricing is approximately $0.14 per million input tokens and $0.56 per million output tokens, making it a cost-efficient option for developers and enterprises.
Read sourceDeepSeek Releases V4.1-Flash Model with Reduced API Prices and Strong Benchmarks
DeepSeek has officially released DeepSeek-V4.1-Flash, the smallest model in its new architecture series, which introduces native multimodal visual understanding capabilities. The model demonstrates strong performance across several benchmarks, including a GPQA Diamond score of 90.9, HLE of 36.8, a Codeforces Rating of 3471, and a Terminal-Bench 2.1 score of 90.6. The release announcement includes complete benchmark figures and details on API switching, allowing users to assess the new model's capabilities and understand the impact of retiring older models on existing integrations and billing. API prices have been reduced accordingly, making the model more accessible for developers and enterprises. This release marks a significant update to DeepSeek's model lineup, emphasizing efficiency and multimodal functionality.