DeepSeek open-sources six Huawei Ascend projects, achieving 99.8% chip peak speed
Chinese AI company DeepSeek has open-sourced six infrastructure components for Huawei's Ascend computing platform, including the TileLang programming language, achieving matrix kernel performance of 99.8% of the chip's theoretical peak speed. The collaboration aims to reduce reliance on Nvidia's CUDA ecosystem amid US export restrictions. DeepSeek founder Liang Wenfeng told investors Huawei could begin delivering training chips this quarter.
Editorial responsibility
- No named human review is recorded for this page.
- Reports are grouped by semantic similarity and deterministic rules. Language models may assist titles, summaries, translation and cross-source analysis; the page reads the event directly, while its address stays stable when the title changes.
- Summary covers the current reports
Cross-source coverage
Common ground
- DeepSeek's work on Huawei's Ascend hardware is a real engineering achievement, showing software optimization can squeeze performance from older chips.
- The US export controls have forced China to build a parallel AI ecosystem, creating a competitor to Nvidia that didn't exist before.
- The hardware fabrication gap—SMIC's older process nodes versus TSMC's latest—is a hard physical constraint that software alone can't fully overcome.
- DeepSeek's stack is currently immature for training frontier models like GPT-4, but it may be sufficient for many inference workloads.
- The 99.8% peak speed figure is a cherry-picked single-kernel benchmark, not representative of real-world production performance.
Points of contention
- Whether this ecosystem is a rational contingency plan or a dangerous walled garden built by a sanctioned company.
- Whether Huawei's legal troubles are proven misconduct or geopolitical smears to justify US dominance.
- Whether the software gap is closing fast enough to make sovereignty considerations outweigh performance differences in the next 3-5 years.
- Whether the US export control strategy is a success that slows China or an own goal that accelerates a rival ecosystem.
- Whether the ethical concerns about transparency and accountability are central or a distraction from engineering realities.
Blind spots
- The debate ignores the role of other countries—like those in Southeast Asia or Europe—that may choose this ecosystem to avoid US dependency.
- No one discussed how China's domestic EUV lithography development is progressing or what realistic timelines exist for closing the foundry gap.
- The long-term cost and energy efficiency of Ascend-based clusters versus Nvidia clusters for large-scale deployment was never compared.
- The potential for this parallel ecosystem to fragment global AI standards and create incompatible technology blocs was overlooked.
WorldAttention’s read
DeepSeek's work on Huawei's Ascend hardware is a real, rational contingency plan for China's AI sector, not a CUDA killer. The software stack shows impressive optimization, but it runs on older transistors from a sanctioned company with legal baggage. The US export controls have achieved their short-term goal of slowing China's access to cutting-edge chips, but they've also forced the creation of a parallel ecosystem that will keep improving. For now, this stack is viable for many inference tasks but not for training frontier models. The real race is against time: can China close the hardware fabrication gap before the software optimization gains run out? Both sides are stuck in ideological battles, missing that the outcome will be decided by engineering progress and geopolitical choices, not moral posturing or nationalist slogans.
Reporting timeline
DeepSeek open-sources six Huawei Ascend projects with matrix kernels at 99.8% peak speed
DeepSeek has open-sourced six projects for Huawei's Ascend AI chips, achieving matrix kernel performance of 99.8% of the chip's theoretical peak speed. The headline project is TileLang, the programming language DeepSeek uses to write its own GPU kernels. Most operators that trained DeepSeek's V4 models are written in TileLang, and each now has an Ascend version, allowing the same source code to compile for Huawei chips instead of Nvidia ones. A second repository, DeepGEMM-Ascend, performs matrix multiplication and reports 431 BF16 TFLOPS against a theoretical ceiling of 432. DeepEP-Ascend, which handles inter-chip data movement, shipped without a license file while the rest use the MIT license, making it currently non-reusable. DeepSeek founder Liang Wenfeng told investors that Huawei could start delivering training chips this quarter, a move that could further reduce reliance on Nvidia.
DeepSeek releases open-source toolkit for Huawei Ascend chips, challenging Nvidia CUDA
DeepSeek has released an open-source toolkit for Huawei's Ascend accelerators, including TileLang support optimized for the Ascend 950 chip. The tools simplify writing the compute kernels behind AI models, with the compiler handling tasks such as scheduling and synchronization. This development could make Huawei hardware more practical for developers building outside Nvidia's CUDA ecosystem. The post, citing Bloomberg, frames this as part of China's broader effort to reduce dependence on Nvidia chips and ramp up its domestic AI chip industry. The author emphasizes that while China is still dependent on Nvidia, it is taking steps to break free from this reliance.
Read sourceDeepSeek partners with Huawei to develop chip programming tools, reducing reliance on Nvidia
Chinese AI company DeepSeek announced on Wednesday a partnership with Huawei Technologies to develop programming tools for Huawei's Ascend chips, marking a significant step in China's efforts to build an alternative to Nvidia's dominant ecosystem. DeepSeek stated on its official WeChat account that it is open-sourcing infrastructure for the Huawei Ascend platform, including a high-level programming language called TileLang, along with related computing and communication libraries. The company said TileLang offers a 'simpler programming model' than Nvidia's CUDA, aiming to improve development efficiency and simplify code logic. This collaboration reflects deepening ties between Chinese tech firms as they seek to reduce dependence on Nvidia's hardware and software ecosystem amid ongoing US export restrictions on advanced semiconductors to China.
Read sourceShow 2 older updatesHide older updates
DeepSeek Open-Sources Ascend Infrastructure Components for Huawei Platform
On September 30, DeepSeek announced the open-sourcing of infrastructure components for Huawei's Ascend computing platform, including the TileLang high-level language compiler, a compute library, and a distributed communication library. These components mirror those previously released for Nvidia's platform. DeepSeek stated that Huawei provided 'unreserved strong support' during the development process, and the two companies collaborated closely on a 128-card supernode solution based on the Ascend 950, achieving deep optimization of computation and communication. The announcement emphasizes continued technological innovation and a commitment to building an open software ecosystem with the community.
DeepSeek Open-Sources Infrastructure Components for Huawei Ascend Computing Platform
DeepSeek, an AI company, has announced the open-sourcing of infrastructure components specifically designed for Huawei's Ascend computing platform. The released components include TileLang, a compilation tool; DeepGEMM; DeepEP; TileKernels; FlashMLA; and DeepSelect. These components correspond to previously open-sourced components for the Nvidia platform, indicating an effort to expand software ecosystem support for Huawei's AI hardware. The announcement, made via DeepSeek's official WeChat public account, provides functional descriptions and open-source links for each component, allowing developers to assess the usability of the Ascend platform's software ecosystem.
Read source