Home/Method notes/Energy per token

Method note · estimate

Energy per token

What it measures

Estimated electricity per token of model output (about three-quarters of a word), measured at the graphics processor (GPU) and scaled to the facility.

Inputs

  • Measured GPU energy per token from published benchmarks
  • The server overhead, and the cooling and building overhead (PUE), from the assumption set stated on the view

Method and weighting

Energy per token starts as joules measured at the GPU for one token of output, converted to watt-hours per thousand tokens.

That GPU figure is multiplied by a server overhead, the extra draw of the machine's own fans, memory and networking beyond the chip itself, to give the energy drawn at the server.

The server figure is then multiplied by PUE to give the energy per token at the whole facility.

Window

As of each benchmark's own date, stated beside it; the assumption set is curated and dated. Data last refreshed on .

Known limits

  • The overheads are single assumptions where real sites vary.
  • A benchmark measures one model on one hardware setup at one batch size.

Used on

Sources: Epoch AI, ML.ENERGY. Each is described on the Data sources page.