NVIDIA Rubin GPU Tops Blackwell by 30x in Agentic AI
NVIDIA has revealed that its upcoming Vera Rubin NVL72 platform achieves up to 30 times the energy efficiency of its Blackwell GB300 NVL72 on a new benchmark for complex, multi-step AI agents.

NVIDIA released performance data demonstrating massive efficiency gains for its next-generation architectures on SemiAnalysis AgentX, an open-source benchmark designed to evaluate agentic AI workloads. Unlike static tests, AgentX replays actual multi-turn coding sessions from Claude Code to capture real-world variables like tool calls and context accumulation. On this benchmark, the upcoming NVIDIA Vera Rubin NVL72 system achieved up to 30 times higher throughput per megawatt than the Blackwell GB300 NVL72 when running the DeepSeek V4-Pro workload at an interactivity target of 160 tokens per second per user.
The benchmark also highlighted the generational leap from Hopper to Blackwell. While a legacy static test on DeepSeek-R1-0528 showed the GB300 NVL72 leading the H200 by up to 40 times more tokens per megawatt, the dynamic AgentX test revealed even more dramatic real-world differences. The GB300 NVL72 delivered up to 15 times higher throughput per megawatt than the older H200 NVL8 when running the 1.6-trillion-parameter DeepSeek V4 Pro. This efficiency translates directly to a 10-fold reduction in token costs for operators. When scaling up to massive mixture-of-experts models like the 2.8-trillion-parameter Kimi K3, the Blackwell system extended its advantage to an 80-fold throughput-per-megawatt gain over the H200 NVL8, while pushing interactivity to 215 tokens per second per user.
These performance leaps are driven by tight integration across NVIDIA's hardware and software ecosystem. Key components include the NVIDIA Dynamo session-aware serving stack, which separates prefill and decode phases, and the high-bandwidth NVLink fabric that coordinates 72 GPUs in a single rack. Software optimizations like DeepGEMM-based kernels, mixed-precision formats such as MXFP4 and MXFP8, and mixture-of-experts runtimes like SGLang, TensorRT-LLM, and vLLM further accelerate execution.
For developers and enterprise practitioners, these advancements address the soaring resource demands of agentic workflows, which consume roughly 15 times more tokens than standard chat and have seen a fourfold growth in average prompt tokens. By converting a fixed power budget into significantly higher interactive throughput, these systems allow organizations to deploy highly complex, multi-step AI agents that can reason and execute tools continuously without incurring prohibitive operational costs or latency bottlenecks.
This is our own summary of reporting by NVIDIA Developer Blog



