Wire flash
TechTongyi Qianwen open-sources Qwen3.8-Flash-Next, early preview of Qwen4 architecture
Editorial responsibility
- No named human review is recorded for this page.
- Source reporting is collected, normalized, translated or condensed automatically when needed.
- Automatically published source-backed update
Tongyi Qianyan (Alibaba's Qwen team) has open-sourced Qwen3.8-Flash-Next, a multimodal Mixture-of-Experts (MoE) model that serves as an early preview of the upcoming Qwen4 architecture. The model features four key upgrades, including GDN + QSA hybrid attention mechanisms, totaling 125 billion parameters with only 6 billion activated per token. Its training cost is approximately one-ninth that of Qwen3.7-Plus, while delivering enhanced performance in coding and office productivity tasks. The early release of Qwen4 architecture weights is significant because QSA sparse attention and N-gram lookup table parameters simultaneously reduce long-context costs and expand capacity, providing early samples for evaluating the next-generation model.
Source report
Tongyi Qianwen has released Qwen3.8-Flash-Next, a multimodal Mixture-of-Experts (MoE) model that serves as an early preview of the Qwen4 architecture.
Key Upgrades
The model introduces four major enhancements, including:
- GDN + QSA hybrid attention mechanism
- Total of 125 billion parameters, with 6 billion activated per token
- Training cost approximately 1/9 that of Qwen3.7-Plus
- Improved performance in coding and office tasks
Significance of the Early Release
The release of Qwen4 architecture weights is critical because:
- QSA sparse attention and N-gram lookup table parameters simultaneously reduce long-context costs and expand capacity
- Provides early samples for evaluating the next-generation model
Source
aihotNeutral / independent
Part of this Story
Alibaba Open-Sources Qwen3.8-Flash, Previewing Efficient Qwen4 Architecture