share_log

The first Ascend 910 super node adopting NPO architecture has arrived! Huawei’s new AI chip is set for an early launch.

cls.cn ·  Sep 17 13:32

① Wang Tao announced that the R&D progress of the Ascend 960 chip has exceeded expectations, achieving a doubling in performance. Specifically, the Ascend 960DT is scheduled for launch in Q1 2027, three quarters ahead of the original plan. ② Huawei released the industry's first Ascend 960 super node utilizing NPO technology, upgraded the Kunpeng super node, and launched the OceanStor M900 AI memory storage system.

Shanghai STAR Market Daily, September 17 (Reporter Huang Xinyi): Today, Huawei Connect 2026 kicked off in Shanghai. Huawei released the industry's first Ascend 960 super node utilizing NPO technology, upgraded the Kunpeng super node, and launched the OceanStor M900 AI memory storage system. Leveraging the Lingqu UnifiedBus interconnect, it builds Agentic super node clusters, supporting ultra-large-scale clusters with millions of cards.

The Shanghai STAR Market Daily also learned from the conference that, to date, the global number of Kunpeng developers has exceeded 4.16 million, with over 7,200 global ecosystem partners, supporting more than 560 global open-source projects. The installed base of openEuler has surpassed 20 million units, making it the leading server operating system in China. The CANN community has seen its monthly active developers break through 5,200, with over 40 native training models based on Ascend + CANN, marking it as the only domestic technical route supporting pre-training.

▍ The core of Huawei's AI strategy is computing power

Wang Tao, Vice Chairman and Rotating Chairman of Huawei, stated that the pace of AI transformation exceeds that of any previous technological revolution in history. Today, large model parameters are rapidly approaching 10 trillion, and are projected to exceed 100 trillion by 2030; AI agents can already operate continuously on an hourly basis, and by 2030, they will execute tasks on a monthly cycle. China's daily inference token volume has grown to approximately 500 trillion, and is expected to reach the quadrillion level by 2030. Edge intelligence is also strengthening rapidly, with mobile models evolving from 3B parameters in 2024 to 30B today, and further progressing towards the 100B level in the future. These changes impose higher requirements on the scale, performance, and reliability of AI infrastructure. Only by building more robust AI infrastructure can we establish a solid foundation for the intelligent world.

Wang Tao stated that the core of Huawei's AI strategy is computing power. The company remains committed to hardware monetization, building competitiveness through system architecture innovation centered on "super nodes + clusters," creating a solid computing foundation for China, and offering new choices for the world. It persists in building an open-source and open computing ecosystem, supporting mainstream large models for native training on Ascend, and actively embracing the "hundreds of models and thousands of scenarios." For customers, Huawei provides flexible computing solutions both offline and online to accelerate the intelligent transformation of various industries. For diverse edge scenarios, it promotes AI integration into devices and vehicles, and upgrades lightweight IoT intelligence by providing diverse computing power ranging from large to micro scales, ensuring ubiquitous intelligence. Furthermore, it constructs next-generation communication networks centered on "computing usage," delivering intelligence to every individual, household, and enterprise.

▍ Release of the Industry's First Ascend 960 Super Node Using NPO

To date, over 1,000 Ascend 910C super nodes have been deployed, and the Ascend 950 super node has also achieved large-scale commercial use. A super node is a computing system physically composed of multiple compute nodes tightly connected via high-efficiency interconnect protocols, featuring unified memory addressing across physical nodes and logically exhibiting the characteristics of a "single computer."

Wang Tao stated that super nodes have become an inevitable choice for building ultra-large-scale AI infrastructure. Currently, a 100,000-card cluster has become the standard for training SOTA models with hundreds of trillions of parameters. However, traditional server architectures cause intra-cluster communication to account for over 40% of training time, severely constraining MFU (Model FLOPS Utilization). Simulation results from Huawei's Markov Lab show that a 100,000-card cluster composed of 4K super nodes achieves a 2.75x improvement in MFU compared to a 100,000-card cluster composed of 8-card servers.

At the conference, Wang Tao announced that the R&D progress of the Ascend 960 chip has exceeded expectations, achieving a doubling in performance. Specifically, the Ascend 960DT is scheduled for launch in Q1 2027, three quarters ahead of the original plan; the Ascend 960PR is scheduled for launch in Q3 2027, one quarter ahead of the original plan. Going forward, the company will maintain an annual evolution cycle, with the Ascend 970 planned for 2028 and the Ascend 980 for 2029. Guided by the innovation direction of Tau Law, computing specifications will continue to double, while memory access bandwidth, memory capacity, and interconnect bandwidth will also see significant improvements.

The design philosophy of the Super Node is to achieve collaboration among multiple NPUs through interconnectivity. Huawei has developed a new generation of Near-Packaged Optics (NPO) optical interconnect products—Hi-ONE. Based on Huawei’s proprietary independent innovation technology, Hi-ONE achieves a transmission capacity of 7.2T per engine through balanced design across optical, mechanical, electrical, magnetic, and thermal elements. This is the industry’s first mass-produced NPO product, currently offering the highest transmission capacity in the sector, and remains the only NPO product with an integrated light source.

Leveraging the Ascend 910 chip and Hi-ONE, Huawei has created the industry’s first Super Node utilizing NPO technology—the Ascend 910 Super Node. A single Super Node scales to 4,096 cards, providing up to 8 ExaFLOPS of FP8 computing power and up to 1PB of HBM capacity. The Super Node employs 5,500 Hi-ONE units, replacing the 48,000 800G optical modules previously required. This reduces power consumption by over 550 kilowatts, doubles the system’s mean time between failures (MTBF), and achieves a system availability of 99.8%.

▍Kunpeng Super Node and AI Memory Storage Debut, Building a Million-Card Super Node Cluster

As the parameter scale of State-of-the-Art (SOTA) models reaches the 10-trillion level, their training and inference systems are no longer singular AI servers or individual intelligent computing Super Nodes, but rather complex computing systems. Such systems comprise intelligent computing Super Nodes, general-purpose computing Super Nodes, and an interconnect system featuring peer-to-peer connectivity without protocol conversion. Additionally, PB-level KV cache clusters must be deployed within the inference system.

It is reported that the Kunpeng Super Node has been upgraded. Based on the Lingqu all-optical networking architecture, the Kunpeng Super Node supports up to 4,096 nodes and forms a unified memory pool of 256TB.

Furthermore, Huawei has developed AI Memory Storage—OceanStor M900—based on the Lingqu architecture. Addressing the multi-level KV Cache architecture required by Agents and long-sequence processing, Huawei proposes a multi-layer KV cache architecture for Agent inference systems. It introduces the Huawei L3.5 layer PB-level KV cache—OceanStor M900 cluster—which is based on Lingqu UnifiedBus and supports “single-hop direct connection.” Meanwhile, the M900 utilizes hybrid media combined with optimized retention algorithms, increasing SSD read/write endurance by 16 times, thereby ensuring KV cache hit rates and long-term stability and reliability at the foundational level.

Simultaneously, Huawei introduces a new Agentic Super Node cluster to accelerate the training and inference of 10-trillion-parameter large models. Based on the Lingqu UnifiedBus interconnect, various interconnect protocols are unified into the Lingqu protocol, significantly reducing protocol conversion overhead. This enables peer-to-peer interconnectivity among subsystems such as Ascend Super Nodes, Kunpeng Super Nodes, and KV cache clusters. It provides a multi-tiered, high-bandwidth, large-capacity storage system where all types of KV caches are accessible via single-hop direct connections. Through a two-layer CLOS four-plane network topology, the maximum cluster scale can reach 512,000 cards. Combined with multi-rail topology technology, it can support an Ascend Super Node cluster of up to one million cards.

The translation is provided by third-party software.


The above content is for informational or educational purposes only and does not constitute any investment advice related to EleBank. Although we strive to ensure the truthfulness, accuracy, and originality of all such content, we cannot guarantee it.