DeepSeek Unveils Cost-Efficient Training Methods for V3 Large Model
Chinese AI startup DeepSeek has published a new technical paper detailing the innovative methods behind its DeepSeek-V3 large language model. The report highlights how the model achieves high efficiency in both training and inference using only 2,048 Nvidia H800 GPUs, a fraction of the tens of thousands typically required by competitors. Co-founder Liang Wenfeng is credited as a contributor to the research. The team attributes this significant cost reduction to four primary technological breakthroughs: memory optimization via multi-head latent attention (MLA), which drastically reduces KV cache usage; a Mixture-of-Experts (MoE) architecture that activates only a subset of parameters per pass; FP8 precision training to halve compute demands; and multi-token prediction for faster inference. These innovations reportedly reduce training costs by 90% compared to dense models with minimal accuracy trade-offs. Additionally, the paper outlines five future directions for AI hardware design, emphasizing the need for tighter software-hardware integration to overcome current bottlenecks in memory, computing, and networking. This release underscores DeepSeek's growing capability to rival major global AI developers through superior architectural efficiency rather than brute computational force.
Wire timeline
DeepSeek Unveils Cost-Efficient Training Methods for V3 Large Model
Chinese AI startup DeepSeek has published a new technical paper detailing the innovative methods behind its DeepSeek-V3 large language model. The report highlights how the model achieves high efficiency in both training and inference using only 2,048 Nvidia H800 GPUs, a fraction of the tens of thousands typically required by competitors. Co-founder Liang Wenfeng is credited as a contributor to the research. The team attributes this significant cost reduction to four primary technological breakthroughs: memory optimization via multi-head latent attention (MLA), which drastically reduces KV cache usage; a Mixture-of-Experts (MoE) architecture that activates only a subset of parameters per pass; FP8 precision training to halve compute demands; and multi-token prediction for faster inference. These innovations reportedly reduce training costs by 90% compared to dense models with minimal accuracy trade-offs. Additionally, the paper outlines five future directions for AI hardware design, emphasizing the need for tighter software-hardware integration to overcome current bottlenecks in memory, computing, and networking. This release underscores DeepSeek's growing capability to rival major global AI developers through superior architectural efficiency rather than brute computational force.
TechNode