Techno Time

Huawei Accelerates AI Infrastructure Push With New Ascend Chips and Million-NPU SuperCluster

Tuesday 22 September 2026 15:24
Huawei Accelerates AI Infrastructure Push With New Ascend Chips and Million-NPU SuperCluster

Huawei has laid out an aggressive roadmap for the next phase of AI infrastructure, unveiling plans for million-NPU computing clusters, new Ascend AI processors and a new generation of SuperPoD systems designed to handle increasingly large foundation models and long-running AI agents.

Speaking at HUAWEI CONNECT 2026 in Shanghai, David Wang, Huawei’s Deputy Chairman of the Board and Rotating Chairman, said the rapid rise of agentic AI is creating computing requirements that conventional server architectures will increasingly struggle to meet.

Huawei’s answer is to build what it describes as a “solid silicon foundation” for the agentic era, combining chips, high-speed interconnects, storage, networking and large-scale computing systems rather than treating AI processors as standalone components.

From 10 trillion to 100 trillion parameters

Huawei expects the scale of AI models to rise dramatically over the remainder of the decade.

Foundation models are approaching 10 trillion parameters, according to the company, and could exceed 100 trillion by 2030. At the same time, AI agents are evolving from systems capable of working continuously for hours toward agents that Huawei expects could handle tasks extending over months by the end of the decade.

The amount of inference is also accelerating. Huawei estimates that daily inference consumption in China alone has reached around 500 trillion tokens, with the figure potentially moving into the quintillions by 2030.

The same trend is reaching personal devices. Smartphone AI models have expanded from around 3 billion parameters in 2024 to 30 billion today, with Huawei expecting on-device models eventually to move into the hundreds of billions.

For Wang, these developments mean the next AI infrastructure challenge will not simply be producing faster chips, but building systems capable of connecting enormous amounts of computing power efficiently.

Huawei accelerates Ascend 960 roadmap

Huawei used the event to provide new details about its next generation of Ascend AI processors.

The company said development of the Ascend 960 has progressed faster than originally planned, with the Ascend 960DT scheduled for the first quarter of 2027, three quarters ahead of its previous roadmap.

The Ascend 960PR is expected in the third quarter of 2027, one quarter earlier than initially planned.

Huawei is targeting a one-generation-per-year development cycle, with the Ascend 970 planned for 2028 and Ascend 980 for 2029.

Wang said the roadmap is intended to deliver improvements not only in raw computing performance but also in memory bandwidth and capacity and interconnect performance — areas becoming increasingly important as AI systems scale.

Why Huawei is betting on SuperPoDs

A major part of Huawei’s strategy revolves around its SuperPoD architecture, which tightly connects multiple computing nodes so they can operate more like a single logical computer.

Huawei said more than 1,000 Atlas 900 A3 SuperPoDs have already been deployed, while its Atlas 950 SuperPoD is entering large-scale commercial use.

The architecture is designed to address a problem that becomes more severe as AI clusters grow: processors can spend a significant amount of time communicating with one another rather than performing model calculations.

Huawei said communications inside conventional 100,000-NPU clusters can consume more than 40% of total model-training time.

Simulation results from Huawei’s Markov Lab indicate that a 100,000-NPU cluster built from 4,000-NPU SuperPoDs could deliver 2.75 times higher Model FLOPs Utilization (MFU) than a comparable cluster constructed from conventional eight-NPU servers.

In practical terms, Huawei is arguing that the AI infrastructure race will increasingly be determined by how efficiently thousands of processors work together, rather than simply by the performance of an individual processor.

Atlas 960E scales to 4,096 NPUs

Huawei also introduced the Atlas 960E SuperPoD, which it describes as the industry’s first SuperPoD based on near-packaged optics, or NPO.

A single system can scale to 4,096 NPUs, delivering up to 8 EFLOPS of FP8 computing performance and as much as 1 petabyte of high-bandwidth memory.

A key component is Huawei’s new Hi-ONE optical interconnect engine, designed to provide 7.2 Tbit/s of transmission capacity per engine.

Huawei said an Atlas 960E using 5,500 Hi-ONE units can avoid the need for around 48,000 conventional 800G optical modules that would otherwise be required to connect the processors.

According to the company, that architecture can reduce power consumption by more than 550 kilowatts while delivering system availability of 99.8%.

From thousands to one million AI processors

Huawei’s ambitions extend considerably beyond individual SuperPoDs.

The company has developed what it calls an agentic SuperCluster, combining Ascend AI SuperPoDs, Kunpeng general-purpose computing systems, large-scale storage and high-speed interconnection through its UnifiedBus architecture.

The system can directly interconnect as many as 512,000 NPUs. With a multi-rail topology, Huawei says the architecture can scale to one million NPUs.

That scale is intended for training and inference workloads associated with models containing around 10 trillion parameters and for agentic systems that require large amounts of context to be stored and accessed rapidly.

Huawei also introduced the OceanStor M900, a storage cluster designed specifically for agentic inference and large-scale KV caching.

The company says the system provides petabyte-scale cache capacity while its storage architecture can extend SSD read/write lifespan by as much as 16 times.

Huawei wants an open ecosystem around its hardware

Hardware represents only one side of Huawei’s AI strategy.

The company said its Kunpeng ecosystem has grown to more than 4.16 million developers and 7,200 partners, while supporting over 560 open-source projects globally.

Huawei said its openEuler operating system has exceeded 20 million installations.

Its Ascend software ecosystem is also expanding. External developers now account for 61% of CANN developers, while the community has more than 5,200 monthly active developers.

More than 40 AI models have been natively pre-trained using Ascend and CANN, according to Huawei, while Ascend supports more than 90 major third-party open-source projects, including PyTorch, Triton, vLLM and veRL.

AI moves from data centers to phones, cars and homes

Huawei’s computing strategy does not stop at massive AI clusters.

The company plans four on-device computing platforms covering AI smartphones, AI PCs, vehicles and homes, combining Kirin and Ascend processors with Pangu and third-party AI models.

HarmonyOS is also being repositioned as what Huawei describes as an “Agent OS”, designed around interaction between people and AI agents across multiple devices.

The longer-term idea is to allow computing resources in devices and the cloud to work together, enabling AI services to move across smartphones, workplaces, vehicles and homes.

Networks will form another layer of that architecture. Huawei expects 5G-A, 6G, 10-gigabit optical networks and low-latency transport networks to connect computing resources spanning data centers, edge infrastructure and devices.

Wang argued that computing power has limited value when isolated, making connectivity a critical part of future AI infrastructure.

The announcements at HUAWEI CONNECT 2026 reveal where Huawei believes the next phase of AI competition is heading. The battle is moving beyond individual AI chips toward complete computing systems capable of making hundreds of thousands — and eventually one million — processors operate as a coordinated infrastructure for increasingly powerful AI agents.