Back to the episode map

Evergreen

How to Evaluate AI Energy Use Claims

Evaluate AI energy and emissions claims by checking the unit, denominator, system boundary, workload, time, geography, method, electricity mix, and uncertainty.

Aug 4, 20267 min readBy Dalton Anderson

How to Evaluate an AI Energy Claim

An AI energy claim is interpretable only when it identifies the unit, denominator, system boundary, workload, phase, time, geography, measurement method, electricity mix, and uncertainty. A global data-center forecast cannot establish the electricity used by one prompt. A lower energy cost per task does not prove that total demand will fall.

The first question is not whether the number is large. It is what the number measures.

flowchart TD
    A["Published AI energy or emissions number"] --> B["Identify unit and denominator"]
    B --> C["Draw the system boundary"]
    C --> D["Name workload, phase, time, and place"]
    D --> E["Separate measured values from allocation and models"]
    E --> F["Check electricity mix, efficiency, and uncertainty"]
    F --> G["Compare only compatible quantities"]
    G --> H["State the narrow supported conclusion"]
    H --> I["Preserve what remains unknown"]

Find the unit and denominator

Energy and power are not the same quantity. Kilowatt-hours, megawatt-hours, and terawatt-hours measure energy over a period. Kilowatts, megawatts, and gigawatts measure a rate of power at a moment or under a stated condition.

Emissions are another quantity. They may be expressed as kilograms or tonnes of carbon dioxide equivalent. Converting electricity into emissions requires an emissions factor and a decision about which electricity mix or accounting method applies.

Every useful rate also needs a denominator. A number might describe energy per prompt, per generated token, per image, per training run, per user, per server, per rack, per facility, per year, or per dollar of revenue.

“AI uses 100 units” is incomplete without both parts.

Draw the system boundary

The boundary determines which equipment and activities the number includes.

At the narrowest level, a measurement may cover accelerator energy during model execution. A broader server boundary can include CPUs, memory, networking, and storage. A facility boundary may add cooling, power conversion, lighting, and other overhead. A grid or lifecycle boundary may include generation losses, construction, equipment manufacture, water, and other upstream effects.

No single boundary is automatically correct. It must match the question.

If the claim concerns data-center electricity demand, facility or grid-level accounting may fit. If it concerns the marginal energy of a model response, a global annual data-center total is too broad.

Name the workload and phase

Training and inference are different activities. Training builds or updates a model. Inference uses a trained model to produce an output.

Inference is not one standardized workload. A short classification, long reasoning task, image generation, video generation, agentic workflow, and enterprise search response can use different models, hardware, context lengths, tool calls, retrieval steps, output volumes, batching, and utilization.

The 2026 IEA Key Questions on Energy and AI record notes that energy use per AI task has been declining while the number of users and energy-intensive uses has increased. That distinction is central. Efficiency can improve at the task level while total electricity demand rises because volume and workload intensity rise faster.

Check whether the value was measured

A meter reading, telemetry estimate, cost allocation, engineering model, company disclosure, and scenario forecast are different forms of evidence.

A measured value still needs a method. Ask where the meter sat, how shared infrastructure was allocated, how idle energy was handled, what sampling period was used, and whether the result represents the tested workload.

A modeled value needs assumptions. Ask about hardware efficiency, utilization, model size, task volume, growth, facility overhead, grid constraints, and technology change.

A forecast is conditional. It describes what may happen under assumptions, not a measurement of the future.

Keep time and geography attached

Hardware, models, data-center design, and electricity systems change. A per-task estimate from an old model and accelerator may not describe a current service.

Geography affects the generation mix, grid losses, emissions intensity, water conditions, local capacity, and the consequence of a concentrated load. A global average can hide a severe local constraint.

Record the measurement period, publication date, geography, facility class, and forecast horizon. Do not repeat the number as timeless.

Separate electricity from emissions

Electricity use does not map to one universal carbon result.

A claim about operational emissions needs an electricity-consumption value and an appropriate emissions factor. The factor may represent a location-based grid average, a market-based contractual claim, a marginal generator, or another method.

Lifecycle emissions can include equipment and construction. Avoid combining a lifecycle figure for one system with an operational figure for another.

State which method was used and why it fits the question. If the source does not provide that information, do not invent precision.

Use system-level evidence at the system level

The IEA’s April 2025 Energy and AI executive summary estimates that data centers consumed about 415 terawatt-hours of electricity in 2024, around 1.5 percent of global electricity consumption. Its base case projects about 945 terawatt-hours in 2030.

The same analysis presents substantial uncertainty after 2030. Its 2035 scenarios span roughly 700 to 1,700 terawatt-hours, with a base case around 1,200 terawatt-hours.

The report's detailed energy-demand analysis keeps the base case and alternative scenarios attached to their assumptions. That is the level at which the figures should be compared.

Those figures support a conclusion about the scale, growth, and uncertainty of global data-center electricity demand. They do not reveal how much electricity one Microsoft Copilot request, Slack summary, model response, or episode-production task used.

The IEA’s April 2026 update says data-center electricity demand increased by 17 percent in 2025 and describes tighter grid, equipment, and planning constraints. It also reports continued task-level efficiency improvement. The update reinforces the difference between falling unit energy and rising total demand.

Audit a dramatic claim

Suppose a headline says an AI prompt uses a fixed amount of electricity.

First, identify the prompt type, model, version, input, output, tool calls, hardware, batch size, utilization, data-center overhead, time, and location. Then determine whether the value was metered, allocated, or modeled. Check whether it describes an average, marginal increment, upper bound, or one experiment.

Next, ask whether the service routes different prompts to different models or performs hidden retrieval and tool work. If it does, one number may not generalize even within the same interface.

The defensible conclusion may be narrow: the study estimated a particular workload under specified assumptions. That can still be useful. It should not become a universal energy label for “an AI prompt.”

Compare like with like

When comparing two systems, align the task outcome and system boundary. A shorter answer that fails the task is not an efficiency win. A model that completes the work in one attempt may outperform a lighter model that requires repeated prompts and human correction.

Include the relevant infrastructure on both sides. Do not compare chip-only energy for one system with facility energy for another. Do not compare an annual fleet forecast with a single inference measurement.

For a workplace decision, the useful unit may be energy per completed, reviewed task. That measure is difficult to obtain, but it describes the operational question better than energy per token alone.

Preserve uncertainty

Report a range when the inputs vary. Identify the assumptions that drive it. State which quantities were unavailable and which effects were excluded.

Precision should reflect the evidence. A result derived from several uncertain allocations should not appear with many confident decimal places.

An uncertain estimate is not worthless. It is more trustworthy when its limits remain attached.

Connect the claim to the decision

The final question is what the number should change.

A company may use system-level forecasts for grid planning, facility siting, procurement, resilience, or policy. A product team may use workload measurements to compare models, shorten context, improve batching, change architecture, or avoid unnecessary generation. A buyer may ask a vendor for task-level evidence and data-center accounting.

Each decision requires a matching boundary.

E042 was right to connect AI adoption with electricity demand. The correction is methodological. Infrastructure announcements, global forecasts, facility loads, and individual tasks are related, but they are not interchangeable evidence.

Editorial note

This guide was developed with AI assistance from the immutable E042 transcript and the linked IEA energy, methodology, and evaluation records. Dalton Anderson remains the author. Energy-methodology, unit, source, current-state, and founder review are mandatory before publication. Publication is not authorized.

Sources

Follow the evidence.

  1. slack.com: 28244420881555 Manage access to AI features in Slackslack.com
  2. learn.microsoft.com: recording transcription overviewlearn.microsoft.com
  3. open.spotify.com: 0FyyANPnMYdcc04GiM2OWXopen.spotify.com
  4. daltonanderson.ghost.io: ai in the workplace is copilot and slack ai worth itdaltonanderson.ghost.io
  5. learn.microsoft.com: security microsoft 365 copilotlearn.microsoft.com
  6. iea.org: key questions on energy and aiiea.org
  7. iea.org: data centre electricity use surged in 2025 even with tightening bottlenecks driving a scramble for solutionsiea.org
  8. slack.com: 31377193680019 Use AI to take huddle notes in Slackslack.com
  9. NIST AI Risk Management Frameworknist.gov
  10. iea.org: executive summaryiea.org
  11. slack.com: 115004846068 Slack updates and changesslack.com
  12. learn.microsoft.com: microsoft 365 copilot overviewlearn.microsoft.com
  13. slack.com: 28310650165907 Security for AI features in Slackslack.com
  14. youtu.be: ZMvMBflUd4youtu.be
  15. NIST Generative AI Profilenvlpubs.nist.gov
  16. slack.com: 25076892548883 Guide to AI features in Slackslack.com
  17. daltonanderson.net: ai in the workplace is copilot and slack ai worth itdaltonanderson.net
How to Evaluate AI Energy Use Claims