Vera Rubin NVL72 inference tests show up to 7x better token throughput per MW vs. Blackwell on a 1.6T DeepSeek model, above Huang's 3x claim for 1T-3T LLMs
First reported by Newsletter.semianalysis ·
AI inference costs drop by up to 7x, making complex models run cheaper and faster.
Nvidia's new Vera Rubin NVL72 platform has demonstrated a significant improvement in inference performance for large language models (LLMs). Tests using a 1.6 trillion parameter DeepSeek model showed up to a sevenfold increase in tokens processed per megawatt compared to the previous Blackwell architecture. This performance leap exceeds earlier projections, including an estimate of a threefold improvement for models between 1 and 3 trillion parameters. The NVL72 system, designed for high-throughput AI inference, leverages Nvidia's latest GPU technology and advanced interconnects to achieve these gains.
The reported sevenfold performance increase for LLM inference on Vera Rubin NVL72 versus Blackwell architecture highlights a major architectural advancement in AI hardware efficiency. This suggests Nvidia is rapidly iterating on its AI compute offerings, pushing the boundaries of tokens per watt and potentially lowering the total cost of ownership for large-scale AI deployments. The discrepancy between actual test results and earlier analyst predictions indicates a potential underestimation of Nvidia's engineering and co-design capabilities for specialized inference workloads.
These efficiency gains have broad implications for the AI market, enabling more sophisticated models to be deployed at scale and reducing the energy footprint of AI computation. Companies relying on AI inference will benefit from lower operational expenses and increased processing power, potentially accelerating the adoption of advanced AI applications. Continued advancements in this area will likely intensify competition among hardware providers and influence the strategic planning of AI development teams globally.
AI-written summary. May contain errors.