Tokens per Wh (efficiency)

Measured inference efficiency: output tokens generated per watt-hour of GPU energy, from the ML.ENERGY Leaderboard (minimum-energy serving configuration; hardware and workload in each value's citation). Higher is better. Closed API providers do not publish per-model energy — a missing value means no credible public measurement exists, not zero.

Measured in tokens/Wh · Higher is better · Weight in overall score: 0

Top models by Tokens per Wh (efficiency)

#NameValue
1Llama 4 Maverick 17B InstructMeta7,435
2Qwen3 Coder PlusAlibaba2,903

Related metrics