vix.ing · top · new · best · stats · spec

Architectural Trade-offs in the Energy-Efficient Era: A Comparative Study of power-capping NVIDIA H100 and H200

2026/04/13 by Aditya Ujeniya, Jan Eitzinger, Georg Hager +1 · 2 voices
Computer Science · Engineering · #Big Data and Digital Economy #Interface (matter) #Key (lock) #Low-power high-performance VLSI design #Memory management #Outlier #Parallel Computing and Optimization Techniques #Power (physics) #Power consumption #cs.PF

paper · pdf · doi:10.48550/arxiv.2604.11391

openalex publication_date 2026/04/13 · arxiv published 2026/04/13 · openalex created_date 2026/04/15 · arxiv updated 2026/07/15 · openalex updated_date 2026/07/28

Abstract

Modern NVIDIA GPUs like the H100 (HBM2e) and H200 (HBM3e) share similar compute characteristics but differ significantly in memory interface technology and bandwidth. By isolating memory bandwidth as a key variable, the power distribution between the memory and Streaming Multiprocessors (SM) changes notably between the two architectures. In the era of energy-efficient computing, analyzing how these hardware characteristics impact performance per watt is critical. This study investigates how the H100 and H200 manage memory power consumption at various power-cap levels. By a regression analysis, we study the memory power limit and uncover outliers consuming more memory power. To evaluate efficiency, we employ compute-bound (DGEMM) and memory-bound (TheBandwidthBenchmark) workloads, representing the two extremes of the Roof\-line model. Our observations indicate that across varying power caps, the H100 remains the slightly better choice for strictly compute-bound workloads, whereas the H200 demonstrates superior efficiency for memory-bound applications.

Citations

Discussions

Related