①DeepSeek has open-sourced a complete set of infrastructure components for Huawei's Ascend platform, which correspond one-to-one with the previously open-sourced components for NVIDIA's platform. ②To build a new-generation, independently controllable GPU software ecosystem, the first priority is to develop a general-purpose, easy-to-program high-level language that can also achieve the hardware's performance ceiling. TileLang was born against this historical backdrop.
According to a report by the STAR Market Daily on September 30 (by reporter Huang Xinyi), DeepSeek announced today that it has officially open-sourced a complete set of infrastructure components for Huawei's Ascend platform, including the TileLang programming language, a computing library, and a distributed communication library, all of which correspond one-to-one with its previously open-sourced components for NVIDIA's platform.
The core of this open-source release is the TileLang Ascend version, which encapsulates the underlying instructions of Ascend, enabling developers to program in a simpler high-level language without sacrificing hardware performance. Its goal is to become a "toolkit" that rivals NVIDIA's CUDA language.
DeepSeek stated that, to build a next-generation, independently controllable GPU software ecosystem, the first priority is to develop a general-purpose, easy-to-program high-level language that can fully exploit the hardware's performance ceiling. It is against this historical backdrop that TileLang was born.
On the one hand, compared with NVIDIA's CUDA language, TileLang is easier to program, significantly improving development efficiency and simplifying code logic.
On the other hand, compared with other high-level languages of the same kind, TileLang's programming model can fully leverage the chip's architectural features, enabling it to reach the hardware's performance ceiling.
The TileLang approach was first validated on NVIDIA's mature platform and now supports the implementation of most operators in the training of the DeepSeek V4 series models, serving as a core tool for exploring new AGI paradigms and developing high-performance operators.
▍Both parties will jointly advance the 128‑card super‑node solution based on the Ascend 950.
Based on the disclosed details, this is far from a simple one-way adaptation. The Huawei team provided substantial support, and the two sides jointly advanced a 128‑card super‑node solution based on the Ascend 950, while also conducting deep optimizations for both computing and communication.
The depth of this "core‑module synergy" is key to enabling components to achieve performance that approaches the hardware's theoretical limits. Relevant technological advances have also been open-sourced in Huawei's CANN community, fostering a mutually reinforcing ecosystem.
According to a reporter from the STAR Market Daily, Huawei has provided the DeepSeek team with jointly defined Ascend SuperNode SuperPoD Flex and UBL128 networking solutions, enabling a single-layer, 3.2 Tbps switch‑based scale‑up network for 128 cards and a two‑layer, switch‑based scale‑out network for 256,000 cards. These solutions can meet the demands of ultra‑low‑latency inference and large‑scale training of cutting‑edge foundation models.
Based on the fully interconnected UBL128 supernode, Ascend provides the ASC-COMM high-performance, custom‑defined communication programming library, enabling users to implement high‑performance programming for both standard communication operators and fused communication‑compute operators. The Deepseek team has developed the high‑performance DeepEP communication library, which supports communication operators across EP, CP, PP, FSDP, and other modes, facilitating the scaling of larger models to larger clusters. Measured interconnect bandwidth reaches 375 GB/s for Dispatch and 347 GB/s for Combine, approaching the hardware‑limited peak performance.
To support developers in deploying the DeepSeek model on Ascend 950 and super-node clusters, Huawei has open-sourced the outcomes of this joint innovation within the CANN community, encompassing high‑EP low‑latency inference deployment, single‑card/single‑node deployment, large‑scale training, and ultra‑long‑text KV cache pooling combined with Agentic RL. Notably, for large‑scale inference scenarios, leveraging the EP32 deployment strategy, under offline inference mode, DeepSeek‑V4.1‑Flash achieves a pure‑model performance of TPOT = 5 ms with a throughput of 2,469 tokens/s per card; at TPOT = 10 ms, the throughput rises to 5,102 tokens/s per card.
In addition, DeepSeek has open-sourced five key computational and communication components that cover the core stages of model training: DeepGEMM accelerates general-purpose matrix operations; DeepEP enables efficient large-scale inter-device communication; TileKernels provides the routine vector‑level computations and memory‑access operators required for data processing; FlashMLA offers sparse attention operators to improve the efficiency of long‑context processing; and DeepSelect implements high‑performance data filtering.
In several key test cases, the computational and communication performance of the aforementioned components has approached the hardware's theoretical limits. Multiple critical tests have demonstrated that both the computational and communication capabilities of these components are nearing their hardware‑level ceilings.
▍Reducing the migration costs of training and inference code for large models
The strategic significance of this open-source initiative far outweighs its technical specifics.
Industry insider Wang Minjian told a reporter from the STAR Market Daily that, over the past three years, the real bottleneck for China's AI chips has never been transistors, but rather software. NVIDIA's competitive moat lies not in chip computing power, but in its CUDA software ecosystem—built up over 18 years—comprising millions of developers, countless optimized operators, and a development experience that lets users "write once, run anywhere." Domestic chips have repeatedly come close to matching NVIDIA's performance specs, yet customers invariably hesitate at the final step: the migration costs are simply too high.
In this announcement, DeepSeek stated that every TileLang operator used in its training has a corresponding high-performance implementation on Ascend. In industry terms, this means that training state-of-the-art models at the DeepSeek scale will no longer rely exclusively on NVIDIA at the software-stack level.
Wang Minjian stated that, for domestically developed chips, TileLang's strategic value lies in a shift in direction: rather than trying to replicate the CUDA ecosystem on its home turf, it is better to establish a new standard at a higher level—a hardware-neutral operator language. Operators are written in TileLang, and the backend can target NVIDIA, Ascend, or any AI chip that is willing to implement a compilation backend.
As early as last year, a brokerage research report pointed out that TileLang could address the issue of interface incompatibility among domestic AI chip–based high-performance computing platforms, thereby reducing the costs associated with migrating code for large‑model training and inference.
▍Domestic AI chips are poised to share an operator ecosystem.
According to publicly available information, TileLang is an open-source, high-performance AI operator programming language developed under the leadership of a team from the School of Computer Science at Peking University. It is a domain-specific language (DSL) that was open-sourced in January 2025. Its core design is based on the "tile" (tensor‑tiling) abstraction and employs a declarative syntax similar to Python, enabling developers to express their computational intent in a form closely resembling mathematical formulas. A compiler then automatically handles low-level hardware optimizations, such as loop tiling and memory scheduling, with the goal of lowering the barrier to AI operator development and achieving "write once, run on multiple architectures."
TileLang was deployed in September 2025 for the development of large-scale models such as DeepSeek‑V3.2‑Exp, and has established compatibility partnerships with hardware vendors including Huawei Ascend, Moore Threads, Suaneng, and Moxie.
At the Developer Day of Huawei Connect 2025, Dong Yuqi, a member of the TileLang team, presented how TileLang enabled the development of the FlashAttention operator, reducing the codebase from over 500 lines to just 80 while maintaining performance on par with the official implementation.
Another noteworthy statement in DeepSeek's announcement reads: As an open-source project, the Ascend‑compatible version of TileLang aims to serve as a model for fostering a highly available software ecosystem across a broader range of AI chips.
Wang Minjian believes that this demonstration effect is aimed at a host of domestic chip companies, including Cambrian, Hygon, Moore Threads, and Muxi: rather than each company reinventing the wheel in isolation, they would do better to jointly interface with a neutral language layer like TileLang. Chip manufacturers would only need to provide compilation backends and low-level instruction interfaces—Ascend has already opened up its Ascend C and PTO ISA instruction sets—while the upper‑level operator ecosystem could be reused across platforms. If this approach proves viable, domestic AI chips could, for the first time, boast an operator ecosystem that spans multiple chip architectures, offering a genuine solution to the fragmentation dilemma.
However, while recognizing the strategic value of this open-source initiative, we must also acknowledge the practical challenges. NVIDIA's CUDA ecosystem has been built over more than a decade, boasting millions of developers and an extensive array of tools. The TileLang Ascend edition is an important starting point, but to move from "usable" to "easy to use" and ultimately to "developers' preferred choice," sustained community engagement and continuous iteration will still be required.
Moreover, while software optimization can approach the hardware's theoretical limits, the hardware's inherent computational ceiling remains the foundation. The Ascend chip's capacity for continuous iteration will determine the maximum scale of models and applications that this software ecosystem can ultimately support. Meanwhile, whether DeepSeek's demonstration can spur more model developers to follow suit will be pivotal in determining whether the ecosystem can truly flourish.
Editor/Deng