vix.ing · top · new · best · stats · spec

Energy use of AI inference, efficiency pathways, and test-time scaling

2025/09/24 by Felipe Oviedo, Fiodar Kazhamiaka, Oviedo, Felipe +13 · 2 voices · 2 citations
Materials Science · Computer Science · Engineering · #Machine Learning in Materials Science #Explainable Artificial Intelligence (XAI) #Advanced Memory and Neural Computing

paper · pdf · doi:10.1016/j.joule.2026.102430

Abstract

As artificial intelligence (AI) inference scales to billions of queries, estimates of per-query energy use are increasingly important for capacity planning, efficiency interventions, and policy. Yet many public estimates assume non-production settings, leading to systematic overestimation. We introduce a bottom-up framework estimating inference energy from token throughput, node power, and overhead under large-scale deployment assumptions. For frontier-scale models (>200B parameters) on H100 nodes, we estimate a median energy of 0.31 Wh/query (interquartile range [IQR] 0.16–0.60), indicating that widely cited estimates are overstated by 4–20×. In test-time scaling scenarios 15× longer than typical queries, the median energy rises 13× to 3.91 Wh (IQR 2.15–7.05). Across models, serving systems, and hardware, we estimate 8–20× line-of-sight energy reductions. At data-center scale, serving 1 billion queries/day requires 0.7 GWh; if 10% are long queries, demand rises to 1.7 GWh/day. With efficiency interventions, it falls to 0.8 GWh/day, mitigating the energy impact of test-time scaling.

Citations

Cited by

Discussions

Related