Back Home

AI 基礎設施

NVIDIA DSX MaxLPS Dynamically Allocates Rack Power, Boosting GB200 Inference Performance per Watt by 1.5×

DSX MaxLPS uses real-time telemetry to reallocate unused power. NVIDIA found that the provisioned power for a GB200 NVL72 rack could be reduced from 125 kW to 90 kW while maintaining throughput. The solution also combines workload-specific power profiles with 45°C liquid cooling, but its core Dynamic Power Software remains in Developer Preview.

Strubbl · CC BY-SA 4.0 · Image source
zh-Hant

NVIDIA has disclosed further details about how DSX MaxLPS works. Instead of permanently reserving power for each GPU rack based on its theoretical peak, Dynamic Power Software (DPS) builds a topology spanning power infrastructure, resource groups, racks, nodes, and GPUs. It continuously monitors actual power consumption and reallocates unused power capacity within administrator-defined budgets and policies. The control loop updates GPU power limits, verifies that groups remain within their aggregate budgets, and handles maintenance or emergency load-shedding events on a best-effort basis.

In a representative inference test published by NVIDIA, MaxLPS reduced provisioned rack power from 125 kW to 90 kW while a GB200 NVL72 ran Kimi-K2.5 at FP4 precision. With total facility power held constant, this could theoretically accommodate 39% more racks and improve performance per watt by about 1.5×. In another test using DeepSeek-R1 at FP4 precision, provisioned power fell from 136 kW to 101 kW, potentially allowing about 35% more racks and improving performance per watt by 1.3× to 1.4×. This is not achieved through downclocking alone: APPM applies prevalidated power profiles based on whether workloads involve inference, training, memory-bound processing, or compute-bound processing, while Dynamo can further optimize inference topologies across racks.

MaxLPS also requires support from facility infrastructure. Design targets for Vera Rubin NVL72 include a coolant inlet temperature of 45°C, enabling data centers in suitable climates to reduce their reliance on compressor-based chillers and redirect cooling power to compute. DSX Exchange can connect DPS with building management, power, cooling, and scheduling systems through a NATS event bus, but it is not required for MaxLPS.

Engineering teams should not directly incorporate the claim of “40% more GPUs” into capacity budgets. These figures are based on NVIDIA-selected models, numerical precision, throughput targets, and its own platforms, and they have not yet been independently reproduced across models. Public materials also contain inconsistent Vera Rubin/GB300 labeling for the second test platform. DPS remains available only to approved early-preview customers, and real-world gains will depend on latency, error rates, peak power consumption, redundancy policies, and site cooling conditions. A more reliable adoption approach is to compare static provisioning with DPS-managed results at the same service level before deciding whether to increase rack density.

Sources

  1. Maximizing AI Factory Performance per Watt with NVIDIA DSX MaxLPS
  2. NVIDIA DSX Documentation