by

Vera Rubin makes Blackwell look pricey

Nvidia claims Vera Rubin can churn agentic AI tokens far more cheaply than Blackwell.

Writing in its developer bog Nvidia  said the latest on-silicon agentic AI results were measured by Nvidia using real-world coding agent trajectories. The benchmark was SemiAnalysis AgentX, which tests inference kit on agentic coding workloads.

AgentX runs across models including Kimi K3, MiniMax M3, GLM5.3, Qwen3.5 and DeepSeek V4 Pro. It is designed to capture the messy bits of agentic work, not just a single neat chatbot turn. The benchmark measures end-to-end interactivity, standard interactivity, end-to-end latency and time to first token. In plain English, it checks whether the machine is fast, responsive and not cooking a power station for fun.

Nvidia says a GB300 NVL72 Grace Blackwell server delivers 15 times higher throughput per megawatt than an H200 NVL8 Hopper setup on DeepSeek V4 Pro 1.6T. The same Blackwell platform is claimed to deliver a cost per million tokens that is 10 times lower than Hopper’s. That means operators can run more agents inside the same power and infrastructure budget, assuming the spreadsheet survives contact with reality.

Blackwell looks stronger as models fatten up. With Kimi K3 2.8T, Nvidia claims an 80-times increase in throughput per megawatt compared to Hopper, while sustaining 215 tokens per second per user. Then Vera Rubin turns up and makes Blackwell look a bit last season. Nvidia claims that the Vera Rubin NVL72 delivers 30 times more throughput than the Grace Blackwell NVL72 on DeepSeek V4 Pro 1.6T.

At about 160 tokens a second per user, Vera Rubin strolls past Grace Blackwell and reaches about 280 tokens a second per user. Blackwell, by comparison, peaks below 180 tokens a second per user. The token cost claim is even more aggressive. Nvidia says Vera Rubin delivers a cost per million tokens that is 35 times lower than Blackwell’s in agentic coding workloads.

Nvidia reckons that means NVL72 can keep several agents running continuously at scale across a broad range of jobs. With DSX MaxLPS managing power across GPUs, racks and workloads, AI factories can provision 40 per cent more GPUs inside the same megawatt budget.

These numbers do not yet include Vera CPU performance for tool calling and focus on the Vera Rubin chips rather than the full seven-chip “Extreme Codesign” platform.

Nvidia says production has started across its AI stack, including Vera CPUs, Rubin GPUs, Vera Rubin servers, Groq 3 LPX chips and networking gear.

 

TOPICS:
agentic ai  ·  AgentX  ·  AI inference  ·  blackwell  ·  hopper  ·  Nvidia  ·  NVL72  ·  tokens  ·  vera rubin

Latest articles

Share

Featured articles

Hot topics

No results found.

Latest reviews