Huawei's Ascend 960 supernode deployment shifts AI computing power competition from single-chip to system engineering.
Author | Link
Huawei Connect 2026 opened in Shanghai on September 17.
Huawei's Vice Chairman and Rotating Chairman, Wang Tao, delivered a keynote speech entitled "Intelligent Future: Building the Silicon-Based Soil for the Intelligent World," and launched the Ascend 960 supernode. Huawei claims this is the industry's first supernode to adopt NPO (Near-Package Optical) technology, with 4096 cards per node, providing a maximum of 8E (EFLOPS, approximately 800 quadrillion calculations per second) FP8 computing power and 1PB HBM capacity.
"The speed of AI transformation surpasses any previous technological revolution in history." In his speech, Wang Tao presented a set of data: the number of parameters in large models is rapidly approaching 10 trillion, and will exceed 100 trillion by 2030; intelligent agents are shifting from working continuously on an hourly basis to executing tasks on a monthly basis; the number of daily inference tokens in China has grown to around 500 trillion, and will reach the trillion-level by 2030; mobile terminal models have evolved from 3B in 2024 to 30B today, and will move towards the 100B level.
These sets of data all point in the same direction: the growth in computing power demand is outpacing the improvement in the performance of individual chips. Huawei's solution is "super nodes + clusters," expanding the competition from individual chips to the entire system.
I. Supernodes become consensus
A supernode refers to a computing system in which multiple computing nodes are tightly coupled through an efficient interconnection protocol, share unified memory, and are logically like a single computer. It addresses the overhead of cluster communication: a 100,000-card cluster is already standard for training trillion-level state-of-the-art models; in traditional server architectures, intra-cluster communication accounts for more than 40% of training time.
Simulation results from Huawei's Markov Lab show that a 100,000-card cluster composed of 4K supernodes achieves a 2.75-fold increase in MFU (Model Floating-Point Utilization) compared to a 100,000-card cluster composed of 8-card servers. MFU directly determines effective computing power; for clusters of the same size, different utilization rates result in significant differences in training costs.
More than 1,000 Ascend 910C supernodes have been deployed, and the Ascend 950 supernode is being used on a large scale. Wang Tao stated that supernodes "have become a consensus among industry, academia, and research in the field of AI infrastructure." NVIDIA's GB200 NVL72 connects 72 GPUs to the same NVLink computing domain, expandable to 576 GPUs, representing a mass-produced form of the same approach.
II. Ascend 960 and NPO Optical Interconnection
The key to the 4096-card supernode lies in interconnection. Huawei's solution is NPO (Near-Packaged Optics), which moves the optical engine from the switch panel to near the chip, shortening the electrical signal transmission distance and reducing power consumption. Huawei released its new generation NPO optical interconnect product, Hi-ONE, with a single-engine transmission capacity of 7.2T. It claims to be the industry's first mass-produced NPO product, the industry's largest transmission capacity, and the only NPO product with a built-in light source.
The Ascend 960 supernode replaces 48,000 800G optical modules with 5,500 Hi-ONE modules, reducing power consumption by over 550 kilowatts, doubling system uptime, and achieving 99.8% availability. Huawei has also proposed an NPO standard project to the global optical interconnect standards organization OIF. In the adjacent co-packaged optics (CPO) route, NVIDIA Spectrum-X Photonics, Broadcom, and others are also making moves; TSMC's silicon photonics platform COUPE is scheduled to enter mass production in 2026.
At the chip level, Huawei stated that the development progress of the Ascend 960 is exceeding expectations, with performance doubling. The 960DT is three quarters ahead of schedule, ready in the first quarter of 2027; the 960PR is one quarter ahead of schedule, ready in the third quarter of 2027. The Ascend 970 and 980 will be launched successively in 2028 and 2029, respectively, relying on the τ law to further double the computing power specifications.
IV. Agent-oriented storage and networking
Agent inference relies on multi-level key-value (KV) caches, requiring storage to shift from simply storing data to supporting inference memory. To address this, Huawei released the OceanStor M900, a petabyte-scale KV cache cluster based on UnifiedBus that supports "one-hop direct connection." Employing hybrid media and an optimized retention algorithm, Huawei claims that SSD read/write lifespan can be improved by 16 times.
The Kunpeng supernode has been upgraded simultaneously, supporting up to 4096 nodes based on the Lingqu all-optical network, forming a unified memory pool of 256TB. Huawei claims that the startup efficiency of the 100,000-level agent sandbox is 30 times higher than that of traditional server solutions, and the efficiency of multiplying the 10 billion-dimensional vector retrieval is doubled compared to traditional solutions.
On top of this, the Agentic supernode cluster unifies and interconnects the Ascend supernode, Kunpeng supernode, and KV cache cluster. Through the two-layer CLOS four-plane networking, the maximum cluster size reaches 512,000 cards. Combined with multi-track topology, it can support an Ascend supernode cluster of up to 1 million cards.
On the network side, Wang Tao said in his speech: "Computing power without a network is an information island." Huawei proposed that communication networks should shift from serving people through connectivity to being built around computing, based on 5G-A/6G, 10 Gigabit optical networks and multi-level collaborative low-latency bearer networks, to deliver intelligence from the center to the edge and end-user side.
The four major computing power platforms on the edge cover AI Phones, AI PCs, cars, and homes. HarmonyOS will be reconstructed from the underlying architecture as Agent OS, "bringing intelligence to every person and every space."
V. Ecological and Industrial Significance
The Ascend ecosystem has achieved two key milestones: CANN has entered into regular open-source community operations, with external developers accounting for 61% for the first time, surpassing internal developers, and monthly active developers exceeding 5,200. Huawei claims this is the most active open-source community in China. Ascend has officially become a computing power platform that can be directly installed from the PyTorch website, which Wang Tao calls "the first computing power platform from China." Currently, Kunpeng has over 4.16 million global developers, over 7,200 ecosystem partners, and over 20 million openEuler installations.
This speech released three noteworthy industry signals.
First, the main battleground for computing power competition is shifting from single chips to system architecture.
NVIDIA's NVL72 and Huawei's supernodes follow the same path, trading interconnectivity and architectural efficiency for effective computing power across entire racks and clusters. The criteria for evaluating computing power companies are shifting from peak computing power to effective computing power.
Second, the optical interconnection industry chain is entering the realization period.
With single-chip bandwidth reaching the Tbit level, the effective transmission distance of copper cables has been compressed to the centimeter level. "Copper out of the cabinet and optical in" has become the engineering choice. Optical interconnects have moved from pluggable modules to near the chip, and the two routes of NPO and CPO are being promoted in parallel.
Upstream sectors such as optical engines, silicon photonics, and lasers are experiencing increased demand, and the focus of competition is shifting to packaging yield, maintainability, and the right to set standards.
Third, storage and networking have gained more weight in the cost structure of the Agent era.
The emergence of KV cache clusters and AI memory storage indicates that "memory" and "connectivity" are beginning to occupy a position on par with "computing power".
"Artificial intelligence may be the last technological revolution for human society. The depth, breadth, and speed of the changes it brings are far beyond imagination. No single company can support the entire intelligent world alone." Wang Tao said at the end of his speech that Huawei will "adhere to open source software to unleash the potential of developers; adhere to embracing diverse models to release value in multiple ways; and adhere to win-win cooperation to ensure customer success and partner growth."
It should also be noted that the Ascend 960 series chips will not be ready until 2027 at the earliest, and the goal of a million-card cluster is still in the planning stage; data such as a 2.75-fold increase in MFU and a 16-fold increase in SSD lifespan are from Huawei's self-testing or simulation and have not been verified by third parties; the positioning of Hi-ONE as the "first mass-produced NPO" needs to be observed in comparison with similar products from NVIDIA, Broadcom and other companies.
Risk Warning and DisclaimerInvesting involves risk; please exercise caution. This article does not constitute personal investment advice and does not take into account the specific investment objectives, financial situation, or needs of individual users. Users should consider whether any opinions, views, or conclusions in this article are suitable for their specific circumstances. Any investment decisions made based on this information are at your own risk.