Trending
Gdańsk University of Technology in Poland submits bid to host Gaia AI Factory Vantage, VoltaGrid face lawsuit concerning natural gas powered off-grid data centers in San Antonio, Texas Building resilience at scale: why modern data centers need integrated risk management Parks S/A to invest R$500m in 5MW data center in Cachoeirinha, Brazil AMD Fires Back at Nvidia with Helios AI System, Epyc CPUs T-Mobile CEO: Satellite exclusivity deal with Starlink doesn’t end this year DCD Talks: A new definition of speed to power with Patrick Dillow, Rowan Digital Infrastructure Google, BlackRock, Ford Motors, and Carhartt launch Alliance for America’s Skilled Trades Prysmian, Relativity Networks pair to deploy their highest density hollow core fiber cable so far Sponsored: The key to the data center power problem Sabey starts construction on 120MW data center campus in Oregon Pantheon Atlas secures grid approval for 1GW Croatia data center Empery Digital invests $20m in Cardinal Data Power AMD and Schneider Electric launch reference standards for Helios AI rack Infleqtion to deploy quantum computer at Illinois Quantum & Microelectronics Park

Nvidia’s LPU gamble

Nvidia is getting bigger.. Late last year, the tech giant swallowed chip startup Groq in a complex $20 billion deal. It was not an acquisition, a pesky formality that includes regulatory oversight, but rather a near-total hiring of staff and licensing of technology..

This feature originally appeared in the AI Hardware supplement. Read it for free today.. “It was really quick,” Nvidia’s VP and GM of hyperscale and HPC Ian Buck tells DCD of the deal, announced on Christmas Eve 2025.

Groq CEO Jonathan Ross “got his Nvidia laptop on Christmas Day; I did a kickoff meeting on the 27th, and it’s been amazing since then.”. The Nvidia Groq 3 LPU – Sebastian Moss. This rapid pace has continued: In March, Nvidia announced it would roll out Groq-based chips before the end of the year, alongside new GPUs and CPUs, with more Groq chips coming on an annual cadence..

The speed of the deal has raised eyebrows, with some seeing it as a desperate move from a dominant force looking to maintain its grip on the AI compute market. Nvidia has seen its market cap skyrocket on the argument that the GPU is the ultimate chip for all things accelerated compute; now it has released an alternative – and had to search outside of the company to find it..

Buck has a unique response. “In 75 percent of use cases, this chip will add very little or no value,” he says, holding up a Groq 3, the company’s take on Groq’s hardware. The language processor unit, or LPU, is targeted at ultra-fast, low-latency AI inference in the agentic AI age..

Key to its speed is Static Random Access Memory, or SRAM. “SRAM is the most expensive kind of memory,” Buck says. “But it’s extremely fast.

This is seven times faster than using HBM, which is the fastest memory technology outside of SRAM that we use,” and what is used in its GPUs.. While the HBM is next to the GPUs, the SRAM is on the chip itself, meaning that data does not need to travel back and forth.. However, the challenge with the LPU is that the amount of SRAM is relatively small, at 500MB.

Even when 256 LPUs are combined in a single rack, that’s 128GB of on-chip SRAM – below the 288GB of HBM4 in a single Vera Rubin.. This is where Groq, the company, struggled, but Groq, the technology, will shine, Nvidia argues. Instead of expecting the LPU to compete, the company will treat it as an addition to the GPU, disaggregating the math and shifting the right parts of the workload to the right chip..

The company will only offer the hardware en masse, with the 256 chips combined in an Nvidia Groq 3 LPX rack. Each tray within the rack will house eight Groq 3 LPUs. It is then set to be packaged alongside a Vera Rubin NVL72 rack, with the more compute-heavy and difficult parts of the decode done on the GPUs and the extreme low-latency work shuffled off to the LPUs..

Using a chatbot as an example, Buck explains: “When you start batching up your context, everything you said in the chat before has to actually be fed in and processed along with the model. That was one of the biggest challenges for the [Groq company era] LPU, because they couldn’t fit all that context memory for the full batches that they were processing.

But we can do all of that processing on here,” he says, brandishing a Rubin, “with the HBM memory.”. – Sebastian Moss. The HBM and GPU handle “all the softmax functions, all of the routing stuff. We don’t have to push any KV calculation to the LPU.

The LPU can just be focusing on the weights. We do have to fit all the weights on there, but only one copy.”. The workload split is “not a traditional disaggregation where you have input tokens being processed by GPU and decodes being done by some ASICs,” Buck argues.

“We’re actually merging the whole thing, and that’s leveraging the software that we have invested in with Dynamo, and they’ve been investing in at Groq, so that we can quickly move the tokens between the two processors.”. CEO and founder Jensen Huang had a more succinct explanation, during a brief interaction: “For the main consumption of tokens, Vera Rubin is unbeatable.

Groq won’t change that. However, we believe that there’s a new segment that’s emerging where the models have to be simultaneously large, the context large, and the latency extremely short.. “Groq can deliver on one of those three promises, but not all three.

However, combining Vera Rubin and Groq, we can deliver on all three of those promises, so that we can create this new segment, large model, big context, and super-fast tokens.”. Huang’s rosy portrayal risks masking a broader message behind the comments – GPUs are still the core backbone of the AI boom, but there’s an admission that space exists at the edges of the workload spectrum for alternative, albeit complementary, chips..

Other SRAM-based inference chip companies, from d-Matrix to Cerebras, have jumped on the news, claiming it validates their approach – which often doesn’t include the Nvidia markup.. The Groq 3 LPU is, in many ways, a chip developed by another company. This initial rollout bears the hallmarks of a rapid integration, with a mix of Nvidia technology, Groq licenses, and third-party hardware..

“They were obviously working on V2 of the Groq chip,” Buck says. “We were able to accelerate the V2 to be V3 in our hands. From a silicon standpoint, obviously, it’s very similar to what they had already, but what they didn’t have was a density form factor with a liquid cooling solution.”.

The LPU will slowly become more Nvidia-ified, with the LPU 35 coming next year and the LPU 40 due alongside the Feynman GPU line. That latter chip will drop Groq’s custom interconnect in favor of Nvidia’s popular NV Link and add NVFP4 support.. The CPU inside the first-gen LPX does not even come from Nvidia; while the company declined to disclose the hardware, racks seen by DCD appear to include AMD Epyc chips..

At some point, it is likely that Nvidia will shift that to its own CPU line, set to be upgraded this year with the Vera CPU.. Like the LPU, it is pitched at the agentic age.. – Sebastian Moss. “The future is kind of starting to look a little more agentic, where you have AIs talking to AIs,” Buck says.

“LLMs are AIs talking to humans, we just took humans out of the loop, so it’s basically however fast this agent can talk to that agent. It’s not gated by human perception.”. Nvidia entered the data center CPU landscape in 2023 with the Grace chip, based on Arm Neoverse cores.

The chip, first announced back in 2021, predated most of the AI boom, let alone the current agentic fever.. “It turns out the workload that you need to connect to a GPU is very similar in nature to what agentic needs,” Buck says. “There’s upfront KV processing, data processing, and of course, having fast cores, because we didn’t want these GPUs to be slow.”.

He points to Amdahl’s Law, where the theoretical maximum speed of a system is limited by the non-parallelizable portion of the task. Grace was developed simply to ensure that the Blackwell GPU hit its potential.. “We have to make sure we always want to have the fastest CPU single-threaded performance to keep the GPU busy.”.

CEO Huang adds: “AI is not about cores, it’s about work. When you have $50 billion of GPUs sitting there, you surely are not going to keep those GPUs idle while $1 billion of CPUs are doing their work.. “You’ve got to get through the work of the CPUs as quickly as possible, so that the $50bn worth of GPUs are never idle.

It means our [CPU requirements] are fundamentally different. And as a result, we came up with very different CPUs.”. But what happened was that customers began using Grace CPUs for more, including as an in-memory KV cache manager across the entire compute fleet. “So, for example, for inference, your query comes in, your KV memory has been sharded across multiple nodes so they can schedule which GPU to go serve your next LLM query,” Buck says. “And then as you update, they maintain that live database and combine all those KV caches.”.

Meta was the first company to publicly announce that it was using both Grace attached to a GPU and as a standalone, but DCD understands there are other companies also deploying standalone Nvidia CPUs.. With the upcoming Vera, Nvidia is more aggressively targeting this market. “This is our next multi-billion dollar business,” Buck says.. “Now I can go to doing agentic work on the CPU,” he says.

“That’s why you’re seeing this explosion in CPU.”. Despite the explosion, which has helped boost the share price of Intel and AMD, Buck calls CPUs “a neglected market.”. He adds: “It got boring because people were simply building another CPU that had more cores and targeting dollar per core.

You never hear about performance. You always hear about dollar per core, dollar per core. As a result, the engineering of those data center CPUs shifted to a place where they just didn’t value performance improvement.”.

The opposite happened in the PC world, where all the focus was on performance and just a few cores. “Unfortunately, it’s not gonna work for agents where you want lots of cores, and you want them all to be fast.”. Vera has 88 cores, notably fewer than a lot of its data center rivals.

“Fewer cores, but the cores are faster, and the fabric that connects all these cores is much stronger,” Buck contests. “It has three times the memory bandwidth per core than x86, two times the energy efficiency.”. By the end of the year, Nvidia hopes to launch Vera – alongside the new Rubin GPU and Groq LPU..

“We were a GPU company,” Buck recalls. “Then we acquired Mellanox, and we became a GPU plus NV Switch plus Spectrum and Quantum and CX and so forth. And now plus Bluefield, Grace, and LPU..

“You can’t serve the entire AI market without all seven chips, all five different racks. The breadth is expanding, models are only getting bigger, all of the AI different modalities are still happening.”. Huang also looked to the past. “We began with the Riva 128,” he says, conveniently ignoring the less successful NV1 that came before it. “We only had one product, now we’ve got 5050-5060-5090, and we’ve got all kinds of products.

It’s no different than the beginning of the iPhone; now there are so many different versions of the iPhone.. “We have customers with different needs, different price points, etc. So we’ve been building out across this spectrum.”

 

Join the conversation

Your email address will not be published. Required fields are marked *