AI infrastructure
First Rubin NVL72 AgentX Benchmarks Released, but Peak Tokens-per-Watt Gains Still Depend on the Baseline
SemiAnalysis tested the Vera Rubin NVL72 using pre-release software and reported up to roughly seven times the token throughput per megawatt of Blackwell. The results highlight the role of HBM4 and NVLink 6 in long-context MoE inference, but the headline 67-fold price-performance gain incorporates rental pricing and assumptions about specific operating points.

SemiAnalysis published the first AgentX results for the Vera Rubin NVL72 on September 14, reporting that pre-release Rubin hardware and software delivered up to roughly seven times the token throughput per megawatt of Blackwell on its long-context, multi-turn coding-agent workload. When its cloud rental-price survey and specific latency operating points were factored in, the performance-per-dollar gap in individual comparisons widened to as much as 67 times. The latter figure is not a fixed hardware multiplier across workloads and should not be applied directly to procurement models.
AgentX measures more than aggregate offline throughput. It produces performance curves based on per-user token generation speed, time to first token, and multi-turn, long-context requests. This better reflects agent services than tokens/s from a single batch because coding agents repeatedly submit long prompts, tool results, and short outputs. If a system’s time to first token rises sharply under high concurrency, even high aggregate throughput may not make it usable.
The hardware design accounts for some of the gains. NVIDIA’s official specifications show that a single NVL72 rack integrates 72 Rubin GPUs, 36 Vera CPUs, 20.7 TB of HBM4, and up to approximately 1.58 PB/s of total memory bandwidth, while NVLink 6 provides 3.6 TB/s of scale-up bandwidth per GPU. For large Mixture-of-Experts (MoE) models, expert-weight reads, the KV cache, and all-to-all communication often become bottlenecks before raw compute does. Treating the entire rack as a single 72-GPU execution domain can reduce cross-node communication and memory pressure.
However, the software remains at an early stage, and the baseline results may use SGLang, TensorRT-LLM, quantization, and parallelism configurations from different points in time. SemiAnalysis’s dollar-based metric also relies on its own survey of three-year rental pricing, while tokens per watt does not account for answer quality or the number of tokens required to complete a task. Engineering teams should wait for raw latency, power, and command-line records from the public workflow and compare results using the same model, precision, context length, and quality-of-service targets. The size of Rubin’s lead will become more meaningful for procurement once AMD MI455X, Google TPU, and production Rubin software are included in the same benchmark.