Alibaba Open-Sources Qwen3.8-Flash, Previewing Efficient Qwen4 Architecture
Alibaba’s Qwen team released Qwen3.8-Flash-Next, an open-weight multimodal Mixture-of-Experts model previewing the Qwen4 architecture. With 125 billion total parameters (6 billion activated per token), it achieves training costs one-ninth of Qwen3.7-Plus while outperforming it in coding and office tasks. The model supports up to 1 million tokens context and is available via OpenRouter and a competitively priced API ($0.16/M input tokens). This release targets cost-sensitive developers and enterprises, expanding Alibaba’s AI ecosystem.
Editorial responsibility
- No named human review is recorded for this page.
- Reports are grouped by semantic similarity and deterministic rules. Language models may assist titles, summaries, translation and cross-source analysis; the page itself is projected from evidence records.
- Source updates only; analysis pending
Analysis pending
Cross-source coverage
Wire timeline
Alibaba's Qwen3.8 Flash Multimodal Model Launches on OpenRouter
Alibaba's Qwen team has released the Qwen3.8 Flash model on the OpenRouter platform. This multimodal reasoning model is designed for a variety of applications including coding assistants, agentic workflows, visual understanding, codebase and document analysis, desktop interaction, charts, and long video processing. The announcement was made via a post on X (formerly Twitter) by the OpenRouter account, highlighting the model's availability and its broad range of capabilities for developers and users.
Alibaba's Qwen Releases Qwen3.8-Flash, Previewing Qwen4 Architecture with Open Weights
Alibaba's Tongyi Qianwen has released Qwen3.8-Flash, an open-weight multimodal Mixture-of-Experts (MoE) model that serves as an early preview of the upcoming Qwen4 architecture. The model features 125 billion total parameters but activates only 6 billion per token, achieving significant efficiency gains. Its training cost is just one-ninth that of the previous Qwen3.7-Plus model, while comprehensively outperforming it on benchmarks, particularly in coding and office tasks. The production API is priced competitively at $0.16 per million input tokens and $0.47 per million output tokens. The model supports a native context length of 262,000 tokens, expandable to 1 million tokens. This release offers a new cost-performance trade-off reference for API call scenarios, targeting cost-sensitive developers and enterprises.
Qwen3.8-Flash-Next-Base Tops 8 of 14 Benchmarks with Only 6B Active Parameters
Alibaba's Qwen team announced the release of Qwen3.8-Flash-Next-Base, a new AI model that achieves top performance on 8 out of 14 major benchmarks, including MMLU-Pro, SuperGPQA, BBH, and GSM8K, despite having only 6 billion active parameters. The model remains competitive with the larger Qwen3.7-Plus-Base on the remaining benchmarks. A key innovation is its 51 billion N-gram embedding parameters, which use deterministic lookups and add no per-token computational cost, enabling efficiency gains.
Show 2 older updatesHide older updates
Alibaba Qwen Announces Qwen3.8-Flash Multimodal MoE Model as Qwen4 Architecture Preview
Alibaba's Qwen team has announced Qwen3.8-Flash, a multimodal Mixture-of-Experts (MoE) model that serves as an early preview of the upcoming Qwen4 architecture. The model is now available as open-weight, with 125 billion parameters plus a 51 billion N-gram component. The production version will be released soon via the QwenCloud API at competitive pricing: $0.16 per million input tokens and $0.47 per million output tokens. This release signals Alibaba's continued advancement in large language model development, offering a powerful yet cost-efficient multimodal AI solution for developers and enterprises.
Qwen3.8-Flash-Next Open Source: Early Preview of Qwen4 Architecture
Tongyi Qianyan (Alibaba's Qwen team) has open-sourced Qwen3.8-Flash-Next, a multimodal Mixture-of-Experts (MoE) model that serves as an early preview of the upcoming Qwen4 architecture. The model features four key upgrades, including GDN + QSA hybrid attention mechanisms, totaling 125 billion parameters with only 6 billion activated per token. Its training cost is approximately one-ninth that of Qwen3.7-Plus, while delivering enhanced performance in coding and office productivity tasks. The early release of Qwen4 architecture weights is significant because QSA sparse attention and N-gram lookup table parameters simultaneously reduce long-context costs and expand capacity, providing early samples for evaluating the next-generation model.