NVIDIA’s latest AI factory design emphasizes three engineered attributes—productivity, durability, and fungibility—by tightly co‑optimising hardware, networking, memory, and the CUDA‑X software stack. For engineers responsible for capacity planning, lifecycle management, or multi‑tenant workloads, the shift means higher token throughput per megawatt, longer hardware ROI, and a single platform that can serve any AI or non‑AI workload without re‑architecting the stack.
Productivity: Tokens per Megawatt and Cost per Token
Power is the primary constraint in large‑scale AI deployments. NVIDIA reports that the Vera Rubin NVL72 system delivers more than 30× the throughput per megawatt compared with the GB300 NVL72, and that token cost can be up to 45× lower on the DeepSeek V4 Pro model. The practical effect is twofold:
- Higher earning capacity: More tokens generated within a fixed power envelope directly increase potential revenue.
- Lower marginal cost: Reducing the cost per million tokens improves margin on each inference or training job.
Engineers should therefore treat tokens / megawatt as a key performance indicator when sizing racks, selecting power distribution, or negotiating capacity contracts.
Durability: Extending the Useful Life of GPU Assets
Historical data shows that NVIDIA GPUs remain economically viable well beyond their nominal depreciation schedules. The A100, launched in 2020, is still in commercial service six years later, with operators extending server depreciation beyond the original six‑year book life. Independent analyses cite useful lives of five to six years for an eight‑GPU H100 system and nine to ten years for the GB300 NVL72, based on resale values. Continuous CUDA kernel optimisations further preserve performance on older silicon.
Implications for practitioners include:
- Plan hardware refresh cycles around actual performance degradation rather than fixed calendar dates.
- Leverage the same CUDA‑X libraries across generations to avoid software rewrites when newer GPUs are introduced.
- Consider extended leasing or resale strategies, as older GPUs retain a significant fraction of their original cost.
Fungibility: A Single Platform for All Workloads
The NVIDIA stack claims to run every AI model type—language, vision, biology, physics, robotics—across all phases from data ingestion to inference, and also supports non‑AI workloads such as scientific simulation and graphics. This breadth is enabled by CUDA’s ability to express parallel math uniformly, allowing a single GPU to execute diverse kernels without hardware changes.
From an operational perspective, this means:
- Infrastructure can be provisioned once and reused for shifting workload mixes, reducing provisioning overhead.
- Security and compliance policies can be applied consistently across AI and non‑AI jobs, simplifying governance.
- Capacity planners can smooth demand spikes by reallocating under‑utilised GPU cycles to alternative workloads.
Related CloudNinjas coverage: AI engineering.
What This Means For Practitioners
Engineers should start measuring token throughput per megawatt as a primary efficiency metric and compare it against cost per token for workload pricing models. Lifecycle strategies must incorporate real‑world depreciation data, favouring software‑only upgrades where possible. Finally, design deployment pipelines around the CUDA‑X library set to maintain workload portability across GPU generations and to exploit the platform’s ability to serve both AI and traditional HPC jobs.


