DeepSeek V4.1 Flash tops coding benchmarks, processes 1 trillion tokens in 24 hours
DeepSeek released V4.1 Flash, a 552-billion-parameter mixture-of-experts AI model that outperforms GPT-6 Sol on coding benchmarks and tops the open-source leaderboard. It features efficient inference with 8 billion active input parameters, native vision support, and costs $0.15 per million input tokens. On OpenRouter, it processed 1 trillion tokens in 24 hours, on pace for the largest 48-hour paid model launch. Optimized versions for local inference on Nvidia Blackwell hardware are also available.
Editorial responsibility
- No named human review is recorded for this page.
- Reports are grouped by semantic similarity and deterministic rules. Language models may assist titles, summaries, translation and cross-source analysis; the page reads the event directly, while its address stays stable when the title changes.
- Summary covers the current reports
Cross-source coverage
Reporting timeline
DeepSeek V4.1 Flash doubles backbone parameters, cuts global KV cache to quarter of V4 Flash
In a post on X, the account Whats_AI reports that DeepSeek V4.1 Flash has nearly twice the backbone parameters of its predecessor V4 Flash, but only about a quarter of its global KV cache. The post notes that this is significant for agents handling long conversations. The author states they have broken down the changes, costs, limits, and results on their own benchmark, providing a link to the analysis. The comparison highlights a trade-off between increased model capacity and reduced memory footprint for long-context tasks, which could impact performance and cost efficiency for AI agents.
Read sourceDeepSeek V4.1 Flash processes 1 trillion tokens in 24 hours on OpenRouter, on pace for record launch
DeepSeek V4.1 Flash has processed 1 trillion tokens in its first 24 hours on the OpenRouter platform, according to a post from OpenRouter. The model is on pace to achieve the largest 48-hour period of any paid model launch, with an estimated 2.8 trillion tokens. Notably, 90% of those tokens were cache reads, priced by the market at approximately $0.006 per million tokens, which is 5 times cheaper than comparable models such as GLM-5.3 Flash. These metrics highlight strong initial adoption and cost efficiency for the new AI model.
Read sourceAtomic Chat squeezes DeepSeek V4.1 Flash 552B MoE into NVFP4 for lower local inference costs
Atomic Chat has released an optimized version of DeepSeek V4.1 Flash, a 552-billion-parameter mixture-of-experts model, compressed into NVFP4 format for more practical local inference on Nvidia Blackwell hardware. The company claims the optimization lowers inference costs without sacrificing core capabilities, retaining a 1 million token context window, native visual understanding, and 98% agreement on agentic dialogue tasks. The weights are available on Hugging Face. This development makes the model a strong fit for local agentic workflows, according to the announcement.
Show 3 older updatesHide older updates
DeepSeek V4.1 Flash tops open-source leaderboard, offers low cost
Thom Wolf announces the release of the DeepSeek V4.1 Flash model, claiming it has returned to the top of the open-source model leaderboard and is extremely cheap. The post highlights the model's efficiency and capability, and Wolf has created a video explaining the forward pass during inference. Links are provided for further reading and to download the model weights. This announcement positions DeepSeek V4.1 Flash as a leading open-source AI model with a focus on cost-effectiveness and performance.
Read sourceDeepSeek V4.1 Flash is a major update, according to analyst post
An X post by analyst Thom_Wolf claims that DeepSeek V4.1 Flash, despite its name suggesting a minor version, represents a major update to the DeepSeek model series. The post provides links to the model weights and an accompanying paper, indicating a significant release in the AI model landscape. The analyst advises not to be distracted by hedging terms like 'flash' or 'minor version' in the name, asserting the update's importance. This announcement is relevant to the AI community, as DeepSeek models are known for their performance and open-weight availability. The post does not specify the exact improvements or benchmarks, but the inclusion of a paper suggests documented advancements. The event is a model release, making it a concrete development in AI research.
Read sourceDeepSeek V4.1 Flash beats GPT-6 Sol on coding benchmarks with 552B parameters
DeepSeek has released V4.1 Flash, a new AI model that reportedly outperforms GPT-6 Sol on several coding and agent benchmarks. The model features 552 billion total parameters, with only 8 billion activating on input and 16 billion on output, enabling efficient inference. It uses just one-quarter the HBM and one-eighth the SSD of DeepSeek V4, includes native vision support, and offers faster inference. DeepSeek V4.1 Flash also beats DeepSeek V4 Pro in speed, cost, and long-horizon coding tasks. Pricing is set at $0.15 per million input tokens and $0.60 per million output tokens, making it a cost-effective option for developers and enterprises.
Read source