Tsinghua University’s KTransformers Enables Local DeepSeek-R1 on Consumer GPUs
The KVCache.AI team from Tsinghua University, in collaboration with APPROACHING.AI, has announced a significant update to the KTransformers open-source project. This development allows users to run the full-parameter DeepSeek-R1 and V3 671B models locally using an NVIDIA RTX 4090D GPU with 24GB VRAM. The solution addresses the high costs and downtime issues associated with cloud-based AI services by leveraging heterogeneous computing, advanced quantization, and sparse attention mechanisms. While inference speeds are lower than enterprise servers, the setup reduces deployment costs by over 95%, bringing the total hardware investment under RMB 70,000. However, the system currently requires specific Intel Xeon CPUs utilizing the AMX instruction set and is optimized primarily for DeepSeek’s Mixture of Experts (MoE) architecture. It supports single-user operations with pre-processing speeds up to 286 tokens per second. This breakthrough highlights a shift towards cost-efficient, localized AI deployment, challenging the dominance of expensive cloud infrastructure and making powerful large language models more accessible to individual developers and smaller organizations despite certain hardware limitations.
Wire timeline
Tsinghua University’s KTransformers Enables Local DeepSeek-R1 on Consumer GPUs
The KVCache.AI team from Tsinghua University, in collaboration with APPROACHING.AI, has announced a significant update to the KTransformers open-source project. This development allows users to run the full-parameter DeepSeek-R1 and V3 671B models locally using an NVIDIA RTX 4090D GPU with 24GB VRAM. The solution addresses the high costs and downtime issues associated with cloud-based AI services by leveraging heterogeneous computing, advanced quantization, and sparse attention mechanisms. While inference speeds are lower than enterprise servers, the setup reduces deployment costs by over 95%, bringing the total hardware investment under RMB 70,000. However, the system currently requires specific Intel Xeon CPUs utilizing the AMX instruction set and is optimized primarily for DeepSeek’s Mixture of Experts (MoE) architecture. It supports single-user operations with pre-processing speeds up to 286 tokens per second. This breakthrough highlights a shift towards cost-efficient, localized AI deployment, challenging the dominance of expensive cloud infrastructure and making powerful large language models more accessible to individual developers and smaller organizations despite certain hardware limitations.
TechNode