Skip to content

Huawei Debuts Atlas 960E Super Node Powered by Hi-ONE Optical Engine

Huawei's Atlas 960E SuperPoD pairs its new Hi-ONE optical engine with 4,096 NPUs and 8 EFLOPS of FP8 compute, introduced at HUAWEI CONNECT 2026 in Shanghai.

Huawei executive presenting the Atlas 960E Super Node racks and specs on stage at Huawei Connect 2026
Hi-ONE · Credit: Huawei

Huawei introduced the Atlas 960E SuperPoD, built around its new Hi-ONE optical engine, at HUAWEI CONNECT 2026 in Shanghai this week. The company calls the machine the Ascend 960 SuperPoD in its Chinese-language materials and pitches it as the industry's first design built on near-packaged optics (NPO) interconnect technology, aimed at training and running trillion-parameter large language models.

The pitch targets a problem that has come to dominate AI cluster design as those models keep growing: getting thousands of processors to behave like a single unit without losing time to network lag. Optical links between chips have become one popular fix for that lag industrywide, and the wager here is that fusing the optics directly onto accelerator packages, rather than bolting on separate pluggable transceivers, buys a real edge in efficiency and uptime once a rack reaches tens of thousands of processors.

Hi-ONE is the piece making that fusion possible. It's described as the first NPO engine to reach mass production anywhere in the sector, rated at 7.2 terabits per second of transmission capacity per unit, and the only such product shipping with an integrated light source rather than depending on an external laser.

David Wang, Huawei's Deputy Chairman of the Board and Rotating Chairman, walked through the hardware in a keynote titled "Advancing the Agentic World, Building a Solid Silicon Foundation." A single unit holds up to 4,096 NPUs, runs at 8 EFLOPS of FP8 throughput and 16 EFLOPS at FP4, and carries as much as a petabyte of HBM. Fitting 5,500 Hi-ONE modules in place of the roughly 48,000 individual 800G transceivers a build that size would traditionally need trims power draw by upward of 550 kilowatts, the company says, while roughly doubling uptime between failures and lifting overall availability to 99.8 percent, per Huawei.

A parallel upgrade rolled the same rack-scale design into the Kunpeng processor family: the TaiShan 950 now links as many as 4,096 nodes over an all-optical fabric, pooling 256 terabytes of shared memory that engineers say speeds up agent-sandbox startup and retrieval work. Sitting next to it, the newly launched OceanStor M900 handles key-value caching at petabyte scale, giving every cache tier a direct single-hop path over that same fabric.

Wired together over LingQu or standard RoCE networking, individual units can grow into far larger deployments; the stated ceiling sits at 512,000 NPUs on a two-tier, four-plane Clos layout, rising to 1 million with an added multi-rail topology. Code got equal billing alongside the silicon: the CANN compute stack shifted to community governance, with backing claimed from more than 90 outside open-source projects, official status as a PyTorch backend, and a developer count on the Kunpeng side put at 4.16 million worldwide. Stripped of the branding, the pitch is straightforward: fewer discrete Hi-ONE modules than a conventional optical build would need, less electricity per rack, and headroom to keep growing toward seven-figure accelerator counts before the wiring between chips becomes the bottleneck instead of the chips themselves.

Share this story

Kenji Sato

Kenji Sato edits techshooked's coverage of artificial intelligence and emerging technology, following the path from research to production systems. His standard is anti-hype: ask what a model actually does, what data trained it, how it fails in practice, and whether a benchmark measures what the marketing says it does.