2017/10/04 by Tomonori Kouya, Kouya, Tomonori
Computer Science · #Advanced Data Storage Technologies #Distributed and Parallel Computing Systems #FOS: Computer and information sciences #G.1.3 #Mathematical Software (cs.MS) #Matrix Theory and Algorithms
paper · pdf · doi:10.48550/arxiv.1710.01839
openalex publication_date 2017/10/04 · openalex created_date 2022/10/06 · openalex updated_date 2026/07/28
Although reliable long precision floating-point arithmetic libraries such as\nQD and MPFR/GMP are necessary to solve ill-conditioned problems in numerical\nsimulation, long precision BLAS-level computation such as matrix multiplication\nhas not been fully optimized because tuning costs are very high compared to\nIEEE float and double precision arithmetic. In this study, we develop a\ntechnique to shorten this tuning time by using prediction of computational\ntimes in several block sizes for the blocking algorithm, and then selecting the\nfastest matrix multiplication method for tuning multiple precision dense real\nmatrix multiplication in various precisions, matrix sizes, and degrees of\nparallelization.\n