At the AI Infra Summit NVIDIA unveiled a set of power‑aware features—DSX MaxLPS dynamic power management, DSX Flex grid‑responsive load control, and tighter hardware integration via NVLink Fusion and custom NVHBM—that collectively raise the tokens‑per‑watt metric for AI factories.
Power‑aware token throughput metric
The industry is moving from raw FLOPS to a measured tokens per megawatt figure. NVIDIA’s DSX MaxLPS claims up to a 1.4× increase in tokens per megawatt by continuously shifting power across GPUs and racks. In a Lambda deployment on Blackwell servers, the same power budget that previously ran 16 full‑power nodes now supports 19 nodes, delivering a 24% rise in token throughput (from ~4 M to ~5 M tokens / s) and a 23% boost in performance‑per‑watt. NVIDIA also projects that next‑generation Vera Rubin NVL72 factories could host up to 40% more GPUs within an unchanged megawatt envelope, translating to roughly 35% higher token throughput without new power infrastructure.
Flexible‑load grid integration
Emerald AI demonstrated a partnership with Silicon Valley Power using the new DSX Flex software. The stack receives real‑time grid signals—load‑shedding requests, demand‑response events, and price cues—and automatically throttles low‑priority jobs while preserving critical workloads. For SREs and capacity planners this introduces a programmable, hierarchical workload policy that can be triggered by external power‑grid events, effectively turning an AI cluster into a demand‑response resource.
Hardware and interconnect updates
Amazon’s Annapurna Labs is co‑developing a custom NVHBM high‑bandwidth memory module for NVIDIA platforms. NVLink Fusion now combines Vera Rubin CPUs with d‑Matrix Raptor XPUs, delivering ultra‑low‑latency inference paths. The broader networking stack includes Spectrum‑X Ethernet, ConnectX SuperNICs, and BlueField DPUs that provide context‑memory storage and infrastructure‑level security. These components are positioned as the physical substrate that enables DSX MaxLPS to shift power at rack scale and to keep latency low when tokens are being generated.
Operational and security considerations
Adopting DSX MaxLPS requires continuous power telemetry across GPUs and racks, as well as integration with existing monitoring pipelines. Workload schedulers must be extended to expose priority tiers that DSX Flex can act upon during grid events. BlueField DPUs, highlighted as part of the stack, introduce a separate processing domain for network and storage security; teams should verify that DPU firmware and policies are aligned with the broader security posture, especially when power‑related control signals traverse the data plane.
Related CloudNinjas coverage: AI engineering.
What This Means For Practitioners
- Incorporate
tokens‑per‑megawattinto capacity‑planning dashboards to surface efficiency gaps early. - Evaluate DSX MaxLPS on existing GPU clusters; the feature works at the rack level and can be enabled without hardware changes on supported Blackwell and Vera Rubin servers.
- Consider DSX Flex if your organization participates in demand‑response programs or operates on a hybrid renewable grid.
- Review interconnect choices—NVLink Fusion, Spectrum‑X, ConnectX SuperNICs—to ensure they meet the latency and bandwidth requirements of token‑intensive workloads.
- Audit BlueField DPU configurations to confirm they enforce the intended isolation between power‑management control paths and data traffic.

