AMD has entered into a technical partnership with rival chip company Cerebras.. The two will develop a disaggregated AI inference solution combining AMD’s new Helios rackscale solution with the Cerebras Wafer-Scale Engine, which features the company’s large AI chip.. The Cerebras Wafer-Scale Engine – Cerebras.
The combined offering is set to be integrated in a single inference workflow, with AMD Helios providing a high-performance, scalable throughput engine for ultra-high throughput, processing prompts and large context windows.. Cerebras will provide ultra-fast, ultra-low latency memory-bandwidth-intensive token generation..
Cerebras plans to deploy AMD Helios systems in its own data centers, with the joint solution expected to become available initially through Cerebras Cloud in the second half of 2026. It will be available more broadly following the cloud rollout.. The increasing demand for ultra-low latency token generation was the reasoning behind Nvidia’s semi-acquisition of Groq, with the company set to deploy the LPX rack featuring Groq LPUs.
Intel has similarly partnered with SambaNova.. “At Cerebras we build the world’s largest and fastest chip,” CEO Andrew Feldman said, detailing a number of large customers. “They deploy us because we’re blisteringly fast.”. He added: “AI has moved from being a novelty to being useful, and in some domains, a necessity.
When it’s a necessity, people want to use it quickly. We saw a partnership where we could extend our footprint in ultra-low latency… It’s really something amazing.”. 16 Feb 2026
