DeepSeek Upgrades V3 Model with Increased Parameters and MIT License
Chinese AI company DeepSeek released an updated version of its DeepSeek-V3 model, designated DeepSeek-V3-0324, on March 24. This new iteration features 685 billion parameters, marking a slight increase from the original model's 671 billion. Although a system card for the update has not yet been published, the company has significantly altered the model's licensing structure by switching to an MIT license. This change aligns the V3 model with the DeepSeek-R1 model, potentially facilitating broader adoption and integration within the open-source community. The original DeepSeek-V3 previously garnered global attention for its exceptional cost-effectiveness and high performance. It outperformed notable open-source competitors like Qwen2.5-72B and Llama-3.1-405B in various benchmarks while achieving results comparable to leading proprietary models such as GPT-4o and Claude-3.5-Sonnet. According to investor High-Flyer Quant, the model was trained at a remarkably low cost of approximately $5.576 million, achieved through optimizations in algorithms, frameworks, and hardware usage. This strategic update underscores DeepSeek's continued commitment to advancing accessible, high-performance artificial intelligence solutions while maintaining competitive efficiency in the rapidly evolving global tech landscape.
Wire timeline
DeepSeek Upgrades V3 Model with Increased Parameters and MIT License
Chinese AI company DeepSeek released an updated version of its DeepSeek-V3 model, designated DeepSeek-V3-0324, on March 24. This new iteration features 685 billion parameters, marking a slight increase from the original model's 671 billion. Although a system card for the update has not yet been published, the company has significantly altered the model's licensing structure by switching to an MIT license. This change aligns the V3 model with the DeepSeek-R1 model, potentially facilitating broader adoption and integration within the open-source community. The original DeepSeek-V3 previously garnered global attention for its exceptional cost-effectiveness and high performance. It outperformed notable open-source competitors like Qwen2.5-72B and Llama-3.1-405B in various benchmarks while achieving results comparable to leading proprietary models such as GPT-4o and Claude-3.5-Sonnet. According to investor High-Flyer Quant, the model was trained at a remarkably low cost of approximately $5.576 million, achieved through optimizations in algorithms, frameworks, and hardware usage. This strategic update underscores DeepSeek's continued commitment to advancing accessible, high-performance artificial intelligence solutions while maintaining competitive efficiency in the rapidly evolving global tech landscape.
TechNode