The AI boom has shifted from brute-force model training into agentic computing, and Nvidia is making sure it owns that layer too.
It claims that infrastructure now lives or dies on memory bandwidth and latency, because agents do not just spit out text; they keep state, juggle tools, and handle long-context workloads.
Nvidia is pitching Blackwell Ultra as the answer, with the GB300 NVL72 rack shown off in fresh benchmarks focused on low latency and long context performance.
Writing in its bog, the company ran Blackwell Ultra through SemiAnalysis’s InferenceMAX and described the results as “astonishing.”
Nvidia’s headline metric is “token/watt”, a number it wants hyperscalers to obsess over while they pour concrete and burn through power budgets.
GB300 NVL72 claims a 50x increase in throughput per megawatt compared with Hopper GPUs, using what it calls the best-deployed state for each architecture.
The company says those gains come from NVLink. Blackwell Ultra scales to a 72-GPU setup, stitching the whole lot into a single NVLink fabric with 130TB/s of connectivity.
That is a very different pitch from Hopper, which Nvidia frames as being stuck with an eight-chip NVLink design, even though many customers did not complain at the time.
Nvidia points to rack design changes and its NVFP4 precision format, arguing this is why the GB300 dominates throughput in inference-heavy scenarios.
Because the industry is currently shouting about “agentic AI”, the company also leans hard on cost per token, which is the bit finance people actually understand.
The company claims a 35x reduction in cost per million tokens with GB300 NVL72, aiming the message squarely at frontier labs and hyperscalers that want more output without the same electricity bill.
It wraps that in the usual scaling-law bravado, saying the pace of improvement is accelerating thanks to its “extreme co-design” approach.
Nvidia admits the Hopper comparison gets messy once you include the different compute nodes and architectural shifts, so it also pitches GB300 against GB200 to look a bit less like it is kicking an old dog.
Across long-context workloads, it claims Blackwell Ultra delivers up to 1.5x lower cost per token and 2x faster attention processing.
The company says context length is a hard constraint for agents, especially when they must hold the full codebase in memory and process tokens to maintain state.
These benchmarks land while Blackwell Ultra is still being integrated by hyperscalers, so Nvidia is treating them as early proof it has kept performance scaling aligned with modern AI workloads.
It cannot resist name-dropping Vera Rubin as the next step, hinting that even more performance is coming, as if anyone doubted the company would sell another rack.
The real message is that Nvidia intends to keep the infrastructure race stitched to its fabric, right down to the tokens, the watts and the invoices.







