DeepSeek releases V4.1-Flash, a 552B MoE model with reduced API prices and strong benchmarks
DeepSeek has released DeepSeek-V4.1-Flash, a 552-billion-parameter Mixture-of-Experts model with a new Encoder-Decoder architecture and native visual understanding. The model activates 8B parameters for input and 16B for output, supports a 1 million token context window, and reduces KV cache to one-quarter of previous versions. API prices have been reduced, with off-peak pricing at $0.14 per million input tokens and $0.56 per million output tokens. The model outperforms competitors like GPT-5.6 Sol on several benchmarks.
Editorial responsibility
- No named human review is recorded for this page.
- Reports are grouped by semantic similarity and deterministic rules. Language models may assist titles, summaries, translation and cross-source analysis; the page reads the event directly, while its address stays stable when the title changes.
- Summary covers the current reports
Cross-source coverage
Reporting timeline
DeepSeek Releases V4.1-Flash with Causal Encoder-Decoder and Visual Understanding
DeepSeek has released V4.1-Flash, the smallest model in its new architecture family, featuring a novel Causal Encoder-Decoder architecture and native visual understanding capabilities. This 552-billion-parameter Mixture-of-Experts (MoE) model activates 8 billion parameters for input processing and 16 billion for output generation. Compared to the previous generation, KV cache requirements have been reduced to one-quarter of HBM and one-eighth of SSD storage, leading to lower API pricing. The article also outlines architectural differences from July's V4-Flash, details parameter and cache metrics, and explains temporary routing arrangements for the V4-Pro API, enabling a quick assessment of the update's substance.
Read sourceDeepSeek AI Releases DeepSeek-V4.1-Flash With 1M Context and Reduced KV Cache
DeepSeek AI has released DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model featuring a 552B-parameter backbone and 196B Engram parameters. The model supports a 1 million token context window and activates 8B parameters during prefill and 16B during decode. A key innovation is the reduction of global KV cache to 890 bytes per token, which is approximately one-quarter of the cache used by DeepSeek-V4-Flash and one-four-hundred-thirty-seventh of DeepSeek-V1. The article details techniques for KV cache compression and pre-pruning architecture, enabling evaluation of long-context deployment costs. This release represents a significant advancement in efficient long-context AI model deployment.
Read sourceDeepSeek releases V4.1 Flash, a 552B parameter MoE model with new architecture
DeepSeek has released V4.1 Flash, a 552 billion parameter Mixture-of-Experts (MoE) model, now available on Hugging Face. The model features a new Encoder-Decoder architecture and has been trained using a novel pre-training method followed by larger-scale reinforcement learning post-training. It is described as the smallest model in DeepSeek's new architecture family and includes native visual understanding capabilities. This release represents a significant update to DeepSeek's model lineup, incorporating architectural innovations and enhanced training techniques.
Read sourceShow 2 older updatesHide older updates
DeepSeek releases V4.1 Flash, a 552B MoE model outperforming GPT-5.6 Sol on benchmarks
DeepSeek has released V4.1 Flash, a new Mixture of Experts (MoE) model with 552 billion total parameters. The model activates only 8 billion parameters for input and 16 billion for output, making it highly efficient. According to benchmark results shared in the announcement, DeepSeek V4.1 Flash outperforms competing models including GPT-5.6 Sol and Opus 5 on several agentic and coding benchmarks. Specific scores include 88.1 on CyberGym versus 84.5 for GPT-5.6 Sol, 54.8 on Automation-Bench versus 45.8, and 74.2 on DeepSWE versus 73.0. The company also announced off-peak API pricing of approximately $0.14 per million tokens for input and $0.56 per million tokens for output, positioning the model as both high-performing and cost-effective for developers and enterprises.
Read sourceDeepSeek Releases V4.1-Flash Model with Reduced API Prices and Strong Benchmarks
DeepSeek has officially released DeepSeek-V4.1-Flash, the smallest model in its new architecture series, which introduces native multimodal visual understanding capabilities. The model demonstrates strong performance across several benchmarks, including a GPQA Diamond score of 90.9, HLE of 36.8, a Codeforces Rating of 3471, and a Terminal-Bench 2.1 score of 90.6. The release announcement includes complete benchmark figures and details on API switching, allowing users to assess the new model's capabilities and understand the impact of retiring older models on existing integrations and billing. API prices have been reduced accordingly, making the model more accessible for developers and enterprises. This release marks a significant update to DeepSeek's model lineup, emphasizing efficiency and multimodal functionality.