2008/07/22 by Kazushige Goto, Robert A. Geijn · 1 citation
Computer Science · #Parallel Computing and Optimization Techniques #Interconnection Networks and Systems #Advanced Data Storage Technologies #Computer science #Parallel computing #Implementation #Matrix multiplication #Computation #Matrix (chemical analysis) #Cache #Computer architecture #Computational science #Algorithm #Programming language
paper · doi:10.1145/1377603.1377607
openalex publication_date 2008/07/22 · openalex created_date 2025/10/10 · openalex updated_date 2026/06/26
A simple but highly effective approach for transforming high-performance implementations on cache-based architectures of matrix-matrix multiplication into implementations of other commonly used matrix-matrix computations (the level-3 BLAS) is presented. Exceptional performance is demonstrated on various architectures.