Moore Threads Deploys DeepSeek Distilled Model on Domestic GPUs
Chinese GPU manufacturer Moore Threads has announced the rapid deployment of inference services for DeepSeek’s distilled large language models on its domestically produced graphics processing units. This strategic move integrates homegrown AI hardware with local large models, aiming to strengthen China’s artificial intelligence ecosystem and reduce reliance on foreign technology. The company utilized a dual-engine approach combining open-source and proprietary technologies to deploy the DeepSeek-R1-Distill-Qwen-7B model via the Ollama framework. Moore Threads highlighted that its self-developed high-performance inference engine enhances computational efficiency through customized operator acceleration and optimized memory management. The deployment demonstrates the versatility of their GPUs and their compatibility with CUDA architecture. This development occurs amidst a surge in DeepSeek’s popularity, with daily active users exceeding 20 million shortly after launch, positioning it as a significant competitor in the global AI landscape. By enabling efficient inference on domestic hardware, Moore Threads aims to empower developers and lay the groundwork for future large-scale model deployments within China’s tech sector.
Wire timeline
Moore Threads Deploys DeepSeek Distilled Model on Domestic GPUs
Chinese GPU manufacturer Moore Threads has announced the rapid deployment of inference services for DeepSeek’s distilled large language models on its domestically produced graphics processing units. This strategic move integrates homegrown AI hardware with local large models, aiming to strengthen China’s artificial intelligence ecosystem and reduce reliance on foreign technology. The company utilized a dual-engine approach combining open-source and proprietary technologies to deploy the DeepSeek-R1-Distill-Qwen-7B model via the Ollama framework. Moore Threads highlighted that its self-developed high-performance inference engine enhances computational efficiency through customized operator acceleration and optimized memory management. The deployment demonstrates the versatility of their GPUs and their compatibility with CUDA architecture. This development occurs amidst a surge in DeepSeek’s popularity, with daily active users exceeding 20 million shortly after launch, positioning it as a significant competitor in the global AI landscape. By enabling efficient inference on domestic hardware, Moore Threads aims to empower developers and lay the groundwork for future large-scale model deployments within China’s tech sector.
TechNode