07What does AI draw from the grid?
The Cost of a Token: the energy behind one AI answer
How much energy does one AI answer take, counted a token (a fragment of text, about three-quarters of a word) at a time?
Measured on the graphics processor (GPU) alone, a generated token costs between 0.05 and 21.9 joules on chat tasks across 25 open models in the ML.ENERGY benchmark, a 468-fold spread that depends on model size, hardware and batching.
Google reports a fleet-wide median of 0.24 Wh for a text prompt, including idle machines and building overhead, by its own account; it is the only whole-system figure among the estimates compared here.
Counting the server around the GPU and the building around the server multiplies the GPU figure by roughly two; the sliders show how much the answer moves with those assumptions.
Benchmark figures are GPU-only under lab conditions. Facility energy depends entirely on the server-overhead and PUE (cooling and building overhead) assumptions shown as sliders. Google's median covers text prompts only.
GPU figures are measured in the ML.ENERGY benchmark (version 3), a public lab test of open models. Server and facility figures are estimates: they multiply the GPU figure by two assumptions, the server overhead and the cooling and building overhead (PUE), which the sliders change.
How these numbers are made: Energy per token.
Get an email when this section changes.