NVIDIA claims 30x agentic gain for Vera Rubin
NVIDIA says its next-generation Vera Rubin NVL72 rack delivers up to 30x the agentic-inference throughput per megawatt of GB300 NVL72, on self-run benchmark results that SemiAnalysis has not yet verified.
NVIDIA published preview results for Vera Rubin NVL72 on SemiAnalysis's AgentX benchmark, reporting up to 30x higher AI-factory throughput per megawatt than GB300 NVL72 on the DeepSeek V4-Pro workload at an interactivity target of 160 output tokens per second per user. The company states the numbers were measured by NVIDIA and are pending SemiAnalysis review.
AgentX is the agentic-coding component of InferenceX, SemiAnalysis's open-source benchmark suite. Rather than fixed prompt-and-response shapes, it replays prerecorded Claude Code sessions turn by turn using the AIPerf client, preserving each session's accumulated context, input/output sequence lengths, reasoning time and tool-call latency — intervals that reproduce the KV-cache pressure of live agent traffic. It varies concurrency and reports tokens per megawatt against four user-experience measures: end-to-end normalized interactivity, standard interactivity, end-to-end latency, and time to first token.
NVIDIA also reports GB300 NVL72 delivering up to 15x the throughput per megawatt of H200 NVL8 on DeepSeek V4 Pro 1.6T under AgentX. On InferenceX's legacy static 8K/1K scenario with DeepSeek-R1-0528, GB300 led H200 by up to 40x — a scenario the post says has been demoted to "maintenance mode" as agentic workloads displace fixed-length serving. NVIDIA cites OpenRouter's State of AI report, covering 100 trillion tokens, that average prompt tokens per request grew roughly fourfold and that a single agentic request consumes 15x the tokens of ordinary chat.
The post gives no absolute tokens-per-megawatt figures, no power-measurement methodology, and no Vera Rubin availability date. All comparisons are vendor-run and "up to"-qualified.