vix.ing · top · new · best · stats · spec

The Design and Performance of Batched BLAS on Modern High-Performance Computing Systems

2017/01/01 by Jack Dongarra, Sven Hammarling, Nicholas J. Higham +3 · 2 citations
Computer Science · #Parallel Computing and Optimization Techniques #Distributed and Parallel Computing Systems #Interconnection Networks and Systems #Computer science #Vectorization (mathematics) #Parallel computing #Supercomputer #Matrix multiplication #Multiplication (music) #Linear algebra #Interface (matter) #Code (set theory) #Double-precision floating-point format #Computational science #Computer architecture #Programming language #Floating point

paper · pdf · doi:10.1016/j.procs.2017.05.138

openalex publication_date 2017/01/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/04

Abstract

A current trend in high-performance computing is to decompose a large linear algebra problem into batches containing thousands of smaller problems, that can be solved independently, before collating the results. To standardize the interface to these routines, the community is developing an extension to the BLAS standard (the batched BLAS), enabling users to perform thousands of small BLAS operations in parallel whilst making efficient use of their hardware. We discuss the benefits and drawbacks of the current batched BLAS proposals and perform a number of experiments, focusing on a general matrix-matrix multiplication (GEMM), to explore their affect on the performance. In particular we analyze the effect of novel data layouts which, for example, interleave the matrices in memory to aid vectorization and prefetching of data. Utilizing these modifications our code outperforms both MKL 1 CuBLAS 2 by up to 6 times on the self-hosted Intel KNL (codenamed Knights Landing) and Kepler GPU architectures, for large numbers of double precision GEMM operations using matrices of size 2 × 2 to 20 × 20.

Citations

Cited by

Related