AMD has officially launched its Helios rackscale system, a wide rack that combines all of the company’s latest AI hardware.. Earlier this week, Microsoft announced plans to deploy the new rack in its data centers. Helios was announced in June last year..
The rack combines AMD Instinct MI455X GPUs, 6th Gen AMD Epyc ‘Venice’ server CPUs, AMD ROCm software, and AMD Pensando networking into a singular rack-scale architecture.. It features 18 compute trays and six switch trays in a 4OU ORW Rack. ORW is a double-width, large-form-factor server frame introduced by the Open Compute Project that spans 1200mm (47.25 inches)..
Each compute tray includes four MI455X GPUs, a single socket Epyc 9006 SP7 Server CPU, and up to three Pensando Vulcano 800 AI NICs per GPU.. The system includes a 50vDC liquid cooled busbar and Rear BlindMate QD liquid-cooling. It is expected to weigh around 5,000 pounds (2,267kg)..
Versus published specifications of rival Nvidia’s Vera Rubin NVL72 rack, the company claims 15 percent more peak FP4 performance, 50 percent more high-bandwidth memory (HBM) capacity, six percent more HBM bandwidth, and 50 percent more scale-out bandwidth. AMD also claims that it will deliver up to 30 percent more tokens per dollar..
A single rack provides up to 2.9 exaflops peak FP4, 1.4 exaflops peak FP8, 31 terabytes of HBM4 memory, and 1.7 petabytes/second of memory bandwidth.. The push into a deeper system-level architecture comes after AMD acquired server maker ZT Systems in 2025 for $4.9bn (to avoid competing with its customers, the company later sold the manufacturing aspect of ZT for $3bn).. “I am very excited to have had the opportunity to come over and work with AMD and the talented architects and engineers here to design this magnificent platform that we call Helios,” former ZT Systems exec and CVP of Platform Architecture, Mark Chubb, said at a press briefing.. “Traditional AI racks are fundamentally discrete GPU servers that are connected through scale-out networks where each server owns its own memory and communicates with other servers over a scale-out network.
Unfortunately, this is typically through multiple network hops, usually introducing latency and congestion.. “Helios changes that model, it’s a rackscale architecture. So instead of nine separate GPU servers that are all connected together through scale-up network, creating a 72 GPU cluster, Helios creates one unified 72 GPU system, all with 31 terabytes of shared HBM memory that is interconnected through a single hop,- UAL or Ethernet fabric – providing approximately 260 terabytes of CLF bandwidth.”.
Helios has now entered full production, with shipments expected by the end of the third quarter.. 13 Apr 2026
