2024/07/01 by Alyson Fox, Fox, Alyson, Peter A. Lindstrom +1
Computer Science · #Advanced Data Storage Technologies #FOS: Mathematics #Numerical Analysis (math.NA) #Numerical Methods and Algorithms #Parallel Computing and Optimization Techniques
paper · pdf · doi:10.48550/arxiv.2407.01826
openalex publication_date 2024/07/01 · openalex created_date 2024/07/06 · openalex updated_date 2026/07/29
The amount of data generated and gathered in scientific simulations and data collection applications is continuously growing, putting mounting pressure on storage and bandwidth concerns. A means of reducing such issues is data compression; however, lossless data compression is typically ineffective when applied to floating-point data. Thus, users tend to apply a lossy data compressor, which allows for small deviations from the original data. It is essential to understand how the error from lossy compression impacts the accuracy of the data analytics. Thus, we must analyze not only the compression properties but the error as well. In this paper, we provide a statistical analysis of the error caused by ZFP compression, a state-of-the-art, lossy compression algorithm explicitly designed for floating-point data. We show that the error is indeed biased and propose simple modifications to the algorithm to neutralize the bias and further reduce the resulting error.