Back to insights
Compute Infrastructure15 August 2026Research note

AI's Energy Constraint Is a Systems Engineering Problem

What global and US electricity estimates imply for AI infrastructure, and how laboratories should measure energy, water, utilization and grid impact.

Institutional analysis947 wordsBy Ram Labs ResearchEvidence reviewed 20 August 2026
Principal finding

AI energy performance cannot be reduced to chip efficiency: useful-work efficiency, facility design, utilization, water, siting, grid constraints and the generation mix determine the system outcome.

415 TWh global data-centre electricity in 2024

IEA estimate, equal to about 1.5% of global electricity consumption; it includes AI and non-AI data-centre loads.

Evidence[1]
945 TWh global demand projected for 2030

IEA base-case projection for all data centres, not a measured value and not attributable entirely to AI.

Evidence[1]
176 TWh US data-centre use in 2023

Lawrence Berkeley National Laboratory estimate, about 4.4% of US electricity.

Evidence[2]
325–580 TWh US 2028 scenario range

LBNL projection corresponding to roughly 6.7%–12% of US electricity; the width expresses genuine uncertainty.

Evidence[2]

Define the system boundary before quoting energy

Energy claims about AI often compare incomparable quantities: a chip's rated power, one model-training run, a data centre, or the global sector. Training and inference have different load shapes, and non-AI services share the same facilities. Electricity consumed at the server is not identical to electricity supplied to the site because cooling, power conversion and networking add overhead. Carbon and water impacts further depend on location, time and technology. Every number should therefore state its boundary, period, geography and whether it is measured, allocated or modelled.

The International Energy Agency estimates that data centres consumed 415 terawatt-hours in 2024, about 1.5% of global electricity. This is a sector total, not an AI-only meter reading. The IEA projects roughly 945 TWh by 2030 in its base case, with AI the most important growth driver alongside other digital services. Reporting the projection as inevitable or assigning all 945 TWh to AI would misstate the analysis.

Evidence[1][6]

Global share can conceal local constraints

A modest global percentage can represent a very large, concentrated load. The IEA reports that nearly half of US data-centre capacity is concentrated in five regional clusters. Individual AI-oriented campuses can resemble power-intensive industrial facilities, but often seek connection on short development timelines. Local transmission, substations, generation adequacy and water availability can therefore constrain projects even when national energy is sufficient. The relevant planning unit is the specific grid node and commissioning schedule.

The US Department of Energy's advisory work notes connection requests for hyperscale facilities of roughly 300 to 1,000 megawatts or more, sometimes with one-to-three-year desired lead times. Those figures describe requested facility scale and timing, not continuous measured consumption. Utilities must test whether requests are firm, coincident and financeable. Developers should provide phased load profiles, ramp rates, reliability needs and flexibility options rather than only a maximum nameplate number.

Evidence[1][3]

The US forecast is a range for a reason

Lawrence Berkeley National Laboratory estimates US data-centre electricity use rose from 58 TWh in 2014 to 176 TWh in 2023, when it represented about 4.4% of US electricity. For 2028, the report presents a wide range of 325 to 580 TWh, or approximately 6.7% to 12% of total US electricity depending on broader demand. The range reflects uncertainty in server shipments, utilization, efficiency, model demand and deployment, not analytical weakness.

Planning should preserve scenarios instead of selecting the largest or smallest figure for advocacy. A laboratory can define a reference case, high-demand case and high-efficiency case, then identify decisions that remain robust across them. Facility commitments with long lead times should use sensitivity analysis for utilization, hardware refresh, electricity price, carbon intensity and curtailment. Forecast error is unavoidable; designs can still be resilient when assumptions and triggers are explicit.

Evidence[2][3]

Measure useful work, not only facility efficiency

Power usage effectiveness compares total facility energy with IT equipment energy and remains useful for infrastructure overhead. It does not reveal whether accelerators are productively utilized or whether a model performs the required task efficiently. Research teams should pair facility metrics with workload measures: joules per verified training result, energy per successful inference, accelerator utilization, memory and network bottlenecks, retries, and the quality threshold achieved. A cheaper but less accurate answer may create more downstream work.

Allocation matters in shared clusters. Teams should distinguish marginal energy from average facility energy, meter at useful temporal resolution, and document how idle capacity and cooling are assigned. Training reports should include hardware, duration, utilization and location; inference reports should include request volume, token or task characteristics, batching and service-level targets. Without these details, a per-query energy number is difficult to reproduce or compare.

Evidence[2][5][6]

Efficiency gains and demand growth coexist

Hardware, quantization, sparsity, model architecture, scheduling and cooling can reduce energy per unit of computation. The DOE reports that computing efficiency has historically improved rapidly and highlights hardware-software co-design. Yet lower unit cost can expand use, and larger models or higher service quality can absorb efficiency gains. Energy strategy should therefore track both intensity and total consumption. An improving joules-per-inference curve does not guarantee that site demand falls.

Infrastructure can also provide flexibility when the workload permits. Some training jobs may shift in time or location, while latency-sensitive inference is less movable. Batteries, onsite generation and thermal storage can support operations, but their value depends on grid rules, emissions and duty cycle. Claims of carbon-free operation should match consumption and generation over an explicit temporal and geographic interval rather than relying only on annual certificate totals.

Evidence[3][5][6]

Make energy an architecture requirement

At project start, set budgets for energy, peak power, water, carbon, latency and cost alongside model quality. Profile the complete stack before buying capacity. Select precision, model size and hardware for the required outcome; use smaller or specialized models where they meet the need. Require experiment tracking to capture energy-relevant configuration, and include infrastructure engineers in model reviews. Siting decisions should consider grid connection, generation mix, climate, water stress, network latency and community impact.

The IEA projects that renewables will meet about half of global growth in data-centre electricity demand to 2035 in its base case, while natural gas and nuclear also contribute. That is a modelled supply pathway, not a guarantee of low-carbon service at every site. A serious research lab should publish measured operational data where feasible, label forecasts, show ranges, and evaluate the resource cost per useful, verified outcome. Energy is not an externality to the compute architecture; it is one of its governing constraints.

Evidence[1][3][4]
Research boundary

Scope and limitations

IEA and LBNL figures model the broader data-centre sector, and AI-specific allocation remains uncertain. Forecasts depend on demand, utilization, technology and electricity-system assumptions. Facility averages can obscure hourly and local effects, while corporate energy claims may use different accounting boundaries. Project decisions require site-level engineering and current utility data.

Evidence base

References

Source review: 20 August 2026. Quantitative values retain their original definitions, periods, and boundaries.

  1. 01
    Energy and AI: Executive summary

    International Energy Agency · 2025

    www.iea.org
  2. 02
    2024 United States Data Center Energy Usage Report

    Lawrence Berkeley National Laboratory · 2024

    datacenters.lbl.gov
  3. 03
    Powering Artificial Intelligence and Data Center Infrastructure

    US Department of Energy Secretary of Energy Advisory Board · 2024

    www.energy.gov
  4. 04
    Energy supply for AI

    International Energy Agency · 2025

    www.iea.org
  5. 05
    AI for Energy

    US Department of Energy · 2024

    www.energy.gov
  6. 06
    Energy and AI

    International Energy Agency · 2025

    www.iea.org