DeepSeek Launches Open-Source AI Model DeepSeek-V3 Rivaling GPT-4o
Chinese AI startup DeepSeek has officially announced the release and open-source launch of its latest large language model, DeepSeek-V3. The announcement was made via a WeChat post, allowing users to interact with the new model on the company's official website. DeepSeek-V3 features a massive architecture with 671 billion total parameters, though only 37 billion are activated during inference, and it was pre-trained on an extensive dataset of 14.8 trillion tokens. Compared to its predecessor, V2.5, the new model offers tripled generation speed, achieving a throughput of 60 tokens per second. While DeepSeek-V3 currently does not support multi-modal input and output, it demonstrates superior capabilities in multilingual processing, particularly excelling in algorithmic code generation and mathematical reasoning. In various benchmark tests, the model has outperformed prominent open-source competitors such as Qwen2.5-72B and Llama-3.1-405B. Furthermore, its performance reportedly matches that of leading proprietary models, including OpenAI's GPT-4o and Anthropic's Claude-3.5-Sonnet, marking a significant advancement in the competitive landscape of artificial intelligence.
Wire timeline
DeepSeek Launches Open-Source AI Model DeepSeek-V3 Rivaling GPT-4o
Chinese AI startup DeepSeek has officially announced the release and open-source launch of its latest large language model, DeepSeek-V3. The announcement was made via a WeChat post, allowing users to interact with the new model on the company's official website. DeepSeek-V3 features a massive architecture with 671 billion total parameters, though only 37 billion are activated during inference, and it was pre-trained on an extensive dataset of 14.8 trillion tokens. Compared to its predecessor, V2.5, the new model offers tripled generation speed, achieving a throughput of 60 tokens per second. While DeepSeek-V3 currently does not support multi-modal input and output, it demonstrates superior capabilities in multilingual processing, particularly excelling in algorithmic code generation and mathematical reasoning. In various benchmark tests, the model has outperformed prominent open-source competitors such as Qwen2.5-72B and Llama-3.1-405B. Furthermore, its performance reportedly matches that of leading proprietary models, including OpenAI's GPT-4o and Anthropic's Claude-3.5-Sonnet, marking a significant advancement in the competitive landscape of artificial intelligence.
TechNode