2017/01/19 by Qiang Wang, Xiaowen Chu, Wang, Qiang +1 · 2 citations
Computer Science · Engineering · #Advanced Data Storage Technologies #FOS: Computer and information sciences #Low-power high-performance VLSI design #Parallel Computing and Optimization Techniques #Performance (cs.PF)
paper · pdf · doi:10.48550/arxiv.1701.05308
openalex publication_date 2017/01/19 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Graphics Processing Units (GPUs) support dynamic voltage and frequency scaling (DVFS) in order to balance computational performance and energy consumption. However, there still lacks simple and accurate performance estimation of a given GPU kernel under different frequency settings on real hardware, which is important to decide best frequency configuration for energy saving. This paper reveals a fine-grained model to estimate the execution time of GPU kernels with both core and memory frequency scaling. Over a 2.5x range of both core and memory frequencies among 12 GPU kernels, our model achieves accurate results (within 3.5%) on real hardware. Compared with the cycle-level simulators, our model only needs some simple micro-benchmark to extract a set of hardware parameters and performance counters of the kernels to produce this high accuracy.