No items found.
No items found.

Why PUE Is Only the Beginning for Efficient AI Factories

PUE for Data Centers and AI Factories
8
Min Read
June 11, 2026
Share Article

AI infrastructure is entering a new phase, where the constraint is no longer just access to GPUs, but the ability to turn every megawatt into productive compute. Power, land, cooling, grid access, and orchestration are becoming the real determinants of scale. That changes the conversation about efficiency. A high-performing AI factory cannot be judged just by how little energy it wastes around the IT load. It should also be judged by how much useful intelligence it produces from the power it consumes.

Power Usage Effectiveness (PUE) remains the clearest measure of data centre discipline. It tells operators, customers, investors, and policymakers how much energy is consumed by supporting infrastructure around the IT load. But PUE measures the facility boundary, not the full production system. A facility can have a strong PUE while GPUs sit idle, training jobs are poorly scheduled, inference workloads produce too few useful tokens per megawatt-hour, or cooling choices reduce electrical overhead while increasing water stress.

The better model is not to abandon PUE. It is to extend it. 

PUE should remain the baseline for facility discipline, especially as regulation makes it more visible and standardized. But AI factories need a broader efficiency model that also measures utilization, workload placement, energy quality, water impact, grid flexibility, and useful compute output. That is the shift from PUE to PUE+.

How the workload has shifted

Conventional cloud data centres were built around mixed, bursty workloads: virtual machines, databases, storage, web services, and enterprise applications. Their efficiency problem was largely about keeping broad pools of CPUs, disks, memory, and networks available across many tenants.

AI factories invert that profile. The dominant load is sustained accelerator compute, surrounded by high-bandwidth networking, fast storage, CPU orchestration, and scheduling systems that keep tightly coupled clusters fed. That is why PUE alone becomes incomplete: the facility may deliver power efficiently, but the economics are determined by whether that power reaches GPUs that are actively training models or serving useful tokens.

For example, a frontier training run is not thousands of small jobs; it is one enormous job, synchronized across thousands of GPUs that compute, communicate, and checkpoint in lockstep. Conventional data center racks drew 5 to 10 kW, but today’s AI factories draw 60 to 140 kW (a single GB200 NVL72 rack pulls 120 kW), beyond what air cooling can physically dissipate. Next-generation platforms are projected to push single-rack consumption past 1 MW

This is the machine PUE now has to describe: denser, hotter, more synchronized, and far less forgiving. The metric still works. It just measures a shrinking share of what determines whether the factory performs. 

Conventional cloud DC AI factory
Conventional cloud rack: 5–15 kW. AI factory rack today: 60–140 kW. Next-gen AI rack (2027–28): 250–1,000 kW.

Bars show typical rack power density ranges (kW per rack), log scale. Sources: TrendForce, Vertiv, Lean Research, 2025–26.

PUE remains a critical baseline for AI infrastructure

PUE is useful because it is simple. It compares total facility energy with IT equipment energy. A facility with a PUE of 1.2 uses 1.2 units of total energy for every 1.0 unit delivered to IT equipment. The remaining 0.2 units are consumed by supporting systems such as cooling, power conversion, pumps, fans, lighting, and other non-IT loads.

That simplicity explains why PUE has lasted. It gives facilities teams a shared language, gives buyers a signal of operational discipline, gives regulators a measurable starting point, and gives investors a way to compare assets beyond broad claims about sustainability or resilience.

In AI infrastructure, this baseline becomes more consequential because the load profile is changing. High-density racks, accelerator clusters, liquid-cooled systems, and tightly coupled networks increase the penalty for inefficiency. Across tens or hundreds of megawatts, a small loss across the electrical path or cooling system becomes a material capacity and cost issue.

The scale of demand makes this difficult to ignore. The International Energy Agency projects global data centre electricity consumption to roughly double from 485 TWh in 2025 to 950 TWh in 2030, with AI-focused data centres tripling electricity consumption over the same period. McKinsey estimates that data centres could require $6.7 trillion in global capital outlays by 2030, including $5.2 trillion for AI processing loads. At that scale, energy overhead is a strategic variable. 

PUE should not be reduced to cooling. Cooling is central, but PUE is shaped by the full mechanical and electrical system: transformers, UPS architecture, redundancy design, distribution losses, airflow containment, pump energy, liquid loops, climate, part-load behavior, and IT demand stability. PUE is the first proof of facility discipline, not the final proof of AI factory efficiency.

Cooling matters, but PUE is not just a cooling story

Cooling deserves serious attention because AI workloads are forcing a new thermal architecture. Direct-to-chip liquid cooling, rear-door heat exchangers, chilled water loops, CDUs, hybrid air-liquid halls, and higher operating temperatures will all shape what future AI factories can support.

But cooling is only one input into PUE. Electrical architecture, redundancy design, utilization, climate, facility maturity, and load consistency also matter. An operator can improve cooling efficiency while losing power elsewhere in the system. It can reduce site-level cooling energy while increasing water intensity. It can design for a low PUE at 100% IT load while operating less efficiently during partial occupancy.

This is why buyers should ask not only for the PUE number, but for the conditions behind it: where it was measured, at what load, across what boundary, in what climate, with what redundancy design, and with what mix of air and liquid cooling. PUE is most useful when it opens better questions. It is least useful when it ends the conversation.

Regulation is making PUE harder to ignore

PUE will remain important not only for technical reasons, but for regulatory ones. Governments are beginning to treat data centres as major energy infrastructure because AI factories influence electricity planning, land use, water strategy, emissions targets, public consent, and national digital capacity.

As a result, PUE is moving from voluntary benchmark to formal reporting and compliance metric. Regulators are beginning with PUE because it is measurable and widely understood, but they are increasingly placing it inside a broader framework that also includes water, renewable energy, waste heat, and grid impact.

In Europe, the revised Energy Efficiency Directive and Delegated Regulation EU/2024/1364 formalizing data centre energy transparency. The framework introduced mandatory reporting requirements for data centres with power demand above 500 kW, created an EU-wide evidence base that included PUE labels and other metrics.

The important point is not only that Europe is asking for data. It is creating comparability.  Starting March 31, 2029, this data will be reviewed every three years to improve sustainability and economic efficiency. Once reporting becomes standardized, facilities can be compared. Once facilities can be compared, ratings become possible. Once ratings exist, performance thresholds can shape procurement, financing, permitting, and public-sector cloud selection.

Singapore is approaching the issue from a different starting point, but the lesson is similar. As a dense digital infrastructure hub with limited land, limited domestic renewable energy, and strong demand for compute, Singapore treats data centre efficiency as a national capacity issue. Its Green Data Centre Roadmap aims to provide at least 300 MW of additional capacity, and guidance points to PUE performance of 1.3 or lower at 100% IT load as an ambition for the next decade. 

In Europe, PUE is becoming part of a reporting and rating system. In Singapore, it is becoming part of the capacity equation for future growth. In both cases, PUE is becoming more formal, visible, and consequential. But regulation also reveals the next problem: even regulators are moving beyond PUE alone.

A strong PUE can still hide a weak AI factory

The limitation of PUE is not that it is wrong. The limitation is that it measures the facility boundary. It tells us how much overhead is required to support IT equipment, but not whether that equipment is doing useful work.

That distinction matters because an AI factory is valuable only when it produces useful compute output. A low-PUE facility can still waste megawatts if accelerators sit idle. A well-cooled cluster can underperform if jobs are fragmented across the wrong racks. A high-density deployment can lose economic value if networking topology limits training efficiency. A facility can look efficient at full load but perform poorly during ramp-up, partial occupancy, or volatile demand.

AI factories need to be measured as production systems that convert constrained inputs into useful output. Their efficiency depends not only on chillers and transformers, but also on utilization, scheduling, model placement, silicon choice, networking, tenancy, power sourcing, and workload economics. A strong PUE is useful, but it is not a complete diagnosis.

GPU utilization is where efficiency becomes productivity

If PUE measures how efficiently power reaches the IT layer, GPU utilization shows whether that power is being turned into productive work. A low-PUE facility with poorly utilized GPUs may look efficient at the building level while wasting the most valuable part of the stack.

GPU utilization is not simply about keeping GPUs busy. It has to mean productive utilization. In training, accelerators should not be stalled by data pipelines, storage bottlenecks, network contention, poor job placement, or inefficient parallelism. In inference, GPUs should serve useful tokens at the right latency, batch size, model configuration, and cost profile.

Utilization also has to be connected to placement. AI workloads are topology-sensitive. A training run spread inefficiently across racks can lose performance to network contention. A multi-tenant inference platform can strand capacity if models are placed without regard to memory footprint, traffic patterns, accelerator type, and thermal zones.

For AI factories, utilization is also a capital efficiency metric. Every idle accelerator represents wasted energy, stranded capex, unused power allocation, and lost revenue capacity. Radiant frames this as the “idle tax”: organizations are not buying intelligence when they buy GPUs; they are buying potential, and utilization determines how much of that potential becomes economic value.

This is the bridge from PUE to PUE+. PUE tells us whether the facility is disciplined with energy overhead. GPU utilization tells us whether the IT load is productive. Tokens per megawatt-hour, training throughput per megawatt, and useful compute per dollar of capex then show whether the AI factory is converting constrained resources into meaningful output.

From PUE to PUE+: the efficiency model AI factories need

The next model should not replace PUE. It should complete it. PUE+ keeps PUE as the baseline for facility discipline, then adds the metrics needed to understand whether an AI factory is productive, sustainable, and investable as a full system.

A practical PUE+ model should answer three questions: how efficiently the facility delivers energy to IT equipment; what resources are consumed to achieve that efficiency; and how much useful AI output is produced from the energy that reaches the IT layer.

Metric What it shows Why it matters
PUE Facility overhead required to support IT load. Baseline measure of facility discipline.
WUE Water consumed per unit of IT energy. Prevents low PUE from masking water-intensive cooling choices.
Energy Reuse Factor Whether waste heat is used productively. Shows whether thermal output is waste or an asset.
Renewable Energy Factor Share of energy from renewable sources. Measures energy quality, not only efficiency.
Carbon intensity per MWh Emissions profile of consumed electricity. Distinguishes efficient power use from low-carbon compute.
GPU utilization Accelerator capacity doing productive work. Reveals whether expensive assets are being wasted.
Tokens per MWh (for inference workloads) Inference output per unit of energy. Measures production efficiency for AI services.
Training throughput per MW (for training workloads) Training progress from constrained power. Links facility design to model development productivity.
Grid flexibility Response to price, carbon, and grid stress. Makes AI factories better energy-system participants.
Useful compute per dollar of capex Productive capacity per unit of investment. Connects engineering efficiency to capital efficiency.

PUE+ is not a single score. It is a layered view of efficiency. It keeps PUE as the baseline, then adds the metrics that reveal whether the AI factory is productive, sustainable, resilient, and economically sound.

The procurement question has to change

For enterprises, governments, telcos, sovereign operators, cloud builders, and investors, the question cannot only be, “What is your PUE?” That question still matters. It should be asked, benchmarked, scrutinized, and reported with clear boundaries. But it should now be followed by a more important question: how much useful AI output can this factory produce from each constrained megawatt?

That question forces the operator to explain utilization, not just cooling. It brings water, carbon, heat reuse, energy quality, workload placement, power availability, software orchestration, and capital efficiency into the same conversation. This is where AI factories will separate from AI-branded data centres.

Radiant’s perspective: PUE is the baseline. PUE+ is the new model.

Radiant’s view is simple: PUE still matters, but AI factories need a bigger efficiency model. Regulation is making PUE more visible and standardized. Power scarcity is making facility efficiency more consequential. AI density is making mechanical and electrical discipline harder to ignore. Capital intensity is making stranded capacity more expensive. But PUE+ is the model that connects those pressures to the real objective: turning constrained power, cooling, land, and silicon into useful compute output.

PUE belongs in that conversation because it is one of the first tests of whether an AI factory is disciplined with power. But the final test is whether the entire system, from powered land and grid access to cooling, compute, networking, scheduling, utilization, and observability, can turn constrained energy into useful intelligence. That is why PUE+ matters. It keeps PUE as the facility baseline, then adds the production metrics that reveal whether the AI factory is actually performing.

That is also why Token /MW, Training Throughput and GPU utilization should sit beside PUE in the AI factory scorecard. PUE shows whether the shell is efficient. Utilization shows whether the engine is producing. Radiant makes that connection operational by treating the GPU estate as a fluid pool of resources, using topology-aware placement to reduce fragmentation, dynamic orchestration and preemption to backfill idle capacity, and GPU fractionalization to run inference workloads that do not need a full GPU. The result is a more productive AI factory, where the same power, cooling, real estate, and silicon produce more useful intelligence.

FAQs

No items found.

How To's

No items found.

Related Articles