Nvidia is pitching its Groq deal as an architectural extension, with chief executive Jensen Huang saying it will act like an accelerator for low-latency decoding.
Groq and Nvidia announced a non-exclusive inference technology licensing agreement on 24 December 2025, alongside the hires of senior Groq executives, and it has repeatedly been described as not being an outright purchase of the company.
On Nvidia’s fourth-quarter earnings call, Nvidia’s chief executive, Jensen Huang, said: “With respect to how we think about Groq and the low latency decoder, I’ve got some great ideas that I’d like to share with you at GTC. And so what we’ll do with Groq is you’ll come to see GTC, but what we’ll do is we’ll extend our architecture with Groq as an accelerator in very much the way that we extended NVIDIA’s architecture with Mellanox.”
Mellanox was about fixing the networking bottleneck and tightening Nvidia’s control of the data centre stack, while Groq is being positioned as a latency play as the industry shifts more attention from training models to running them at scale.
The Groq angle is inference, particularly the decoding step where models generate tokens, which is where responsiveness becomes the product for chat, agents and other interactive workloads.
Nvidia is signalling it wants Groq’s low-latency strengths to sit alongside its GPU platform rather than replace it, and it is using GTC as the venue to explain how that integration is meant to work.







