07What does AI draw from the grid?

The Cost of a Token: the energy behind one AI answer

How much energy does one AI answer take, counted a token (a fragment of text, about three-quarters of a word) at a time?

Measured on the graphics processor (GPU) alone, a generated token costs between 0.05 and 21.9 joules on chat tasks across 25 open models in the ML.ENERGY benchmark, a 468-fold spread that depends on model size, hardware and batching.

Google reports a fleet-wide median of 0.24 Wh for a text prompt, including idle machines and building overhead, by its own account; it is the only whole-system figure among the estimates compared here.

Counting the server around the GPU and the building around the server multiplies the GPU figure by roughly two; the sliders show how much the answer moves with those assumptions.

Benchmark figures are GPU-only under lab conditions. Facility energy depends entirely on the server-overhead and PUE (cooling and building overhead) assumptions shown as sliders. Google's median covers text prompts only.

Published energy per query, and how each was arrived at
Wh per query or prompt; one figure per source
Energy per generated token, by model and GPU

GPU figures are measured in the ML.ENERGY benchmark (version 3), a public lab test of open models. Server and facility figures are estimates: they multiply the GPU figure by two assumptions, the server overhead and the cooling and building overhead (PUE), which the sliders change.

Loading data…

How these numbers are made: Energy per token.

Get an email when this section changes.