2016/03/01 by Marat Dukhan, Dukhan, Marat, Richard Vuduc +3
Computer Science · Engineering · #FOS: Computer and information sciences #FOS: Mathematics #Low-power high-performance VLSI design #Numerical Analysis (math.NA) #Numerical Methods and Algorithms #Parallel Computing and Optimization Techniques #Performance (cs.PF)
paper · pdf · doi:10.48550/arxiv.1603.00491
openalex publication_date 2016/03/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/01
We propose a new instruction (FPADDRE) that computes the round-off error in floating-point addition. We explain how this instruction benefits high-precision arithmetic operations in applications where double precision is not sufficient. Performance estimates on Intel Haswell, Intel Skylake, and AMD Steamroller processors, as well as Intel Knights Corner co-processor, demonstrate that such an instruction would improve the latency of double-double addition by up to 55% and increase double-double addition throughput by up to 103%, with smaller, but non-negligible benefits for double-double multiplication. The new instruction delivers up to 2x speedups on three benchmarks that use high-precision floating-point arithmetic: double-double matrix-matrix multiplication, compensated dot product, and polynomial evaluation via the compensated Horner scheme.