vix.ing · top · new · best · stats · spec

GPU Performance Portability needs Autotuning

2025/04/30 by Burkhard Ringlein, Ringlein, Burkhard, Thomas Parnell +3 · 1 citation
Computer Science · #Parallel Computing and Optimization Techniques #Advanced Neural Network Applications #Advanced Data Storage Technologies

paper · pdf · doi:10.48550/arxiv.2505.03780

Abstract

As LLMs grow in complexity, achieving state-of-the-art performance requires tight co-design across algorithms, software, and hardware. Today's reliance on a single dominant platform limits portability, creates vendor lock-in, and raises barriers for new AI hardware. In this work, we make the case for combining just-in-time (JIT) compilation with comprehensive kernel parameter autotuning to enable portable LLM inference with state-of-the-art performance without code changes. Focusing on performance-critical LLM kernels, we demonstrate that this approach explores up to 15x more kernel parameter configurations, produces significantly more diverse code across multiple dimensions, and even outperforms vendor-optimized implementations by up to 230%, all while reducing kernel code size by 70x and eliminating manual code optimizations. Our results highlight autotuning as a promising path to unlocking model portability across GPU vendors.

Cited by

Related