2021/10/01 by Guin Gilman, Gilman, Guin, Robert J. Walls +1
Computer Science · #Advanced Data Storage Technologies #Artificial intelligence #Computer architecture #Computer science #Concurrency #Deep learning #Distributed #Distributed and Parallel Computing Systems #FOS: Computer and information sciences #Hardware Architecture (cs.AR) #Inference #Machine Learning (cs.LG) #Microarchitecture #Operating system #Parallel #Parallel Computing and Optimization Techniques #Parallel computing #Preemption #Scheduling (production processes) #Stochastic Gradient Optimization Techniques #and Cluster Computing (cs.DC) #cs.AR #cs.DC #cs.LG
paper · pdf · doi:10.48550/arxiv.2110.00459
published in arXiv (Cornell University) (Cornell University) · To Appear in the 39th International Symposium on Computer Performance, Modeling, Measurements and Evaluation (Performance 21)
arxiv created 2021/10/01 · openalex publication_date 2021/10/01 · arxiv updated 2021/10/04 · openalex created_date 2022/10/01 · openalex updated_date 2026/08/05
We investigate the performance of the concurrency mechanisms available on\nNVIDIA's new Ampere GPU microarchitecture under deep learning training and\ninference workloads. In contrast to previous studies that treat the GPU as a\nblack box, we examine scheduling at the microarchitectural level. We find that\nthe lack of fine-grained preemption mechanisms, robust task prioritization\noptions, and contention-aware thread block placement policies limits the\neffectiveness of NVIDIA's concurrency mechanisms. In summary, the sequential\nnature of deep learning workloads and their fluctuating resource requirements\nand kernel runtimes make executing such workloads while maintaining\nconsistently high utilization and low, predictable turnaround times difficult\non current NVIDIA hardware.\n