Tencent Optimizes DeepSeek's DeepEP Framework, Doubling RoCE Performance
Chinese AI startup DeepSeek has credited Tencent with significantly enhancing the performance of its open-source DeepEP communication framework. According to DeepSeek, Tencent’s Starlink Networking team identified and resolved critical bottlenecks, specifically underutilized dual-port NIC bandwidth and CPU control latency. These targeted optimizations resulted in a 100% performance improvement on RoCE networks and a 30% gain on InfiniBand (IB) environments. DeepEP is a specialized communication library designed for Mixture of Experts (MoE) and expert parallelism, supporting high-throughput, low-latency GPU kernels and low-precision computing such as FP8. The enhanced framework has been fully open-sourced on GitHub, where DeepSeek acknowledged the contribution as a huge speedup. Furthermore, the optimized DeepEP has been successfully deployed in the training of Tencent’s Hunyuan large language model, demonstrating its versatility within infrastructure built on Tencent’s Starlink and H20 servers. This collaboration highlights a significant advancement in efficient solutions for AI model training, offering improved capabilities for developers working with complex neural network architectures across different network settings.
Wire timeline
Tencent Optimizes DeepSeek's DeepEP Framework, Doubling RoCE Performance
Chinese AI startup DeepSeek has credited Tencent with significantly enhancing the performance of its open-source DeepEP communication framework. According to DeepSeek, Tencent’s Starlink Networking team identified and resolved critical bottlenecks, specifically underutilized dual-port NIC bandwidth and CPU control latency. These targeted optimizations resulted in a 100% performance improvement on RoCE networks and a 30% gain on InfiniBand (IB) environments. DeepEP is a specialized communication library designed for Mixture of Experts (MoE) and expert parallelism, supporting high-throughput, low-latency GPU kernels and low-precision computing such as FP8. The enhanced framework has been fully open-sourced on GitHub, where DeepSeek acknowledged the contribution as a huge speedup. Furthermore, the optimized DeepEP has been successfully deployed in the training of Tencent’s Hunyuan large language model, demonstrating its versatility within infrastructure built on Tencent’s Starlink and H20 servers. This collaboration highlights a significant advancement in efficient solutions for AI model training, offering improved capabilities for developers working with complex neural network architectures across different network settings.
TechNode