2019/11/11 by Bruno Bauwens, Bauwens, Bruno, Marius Zimand +1
Computer Science · Mathematics · #68Q30 #94A15 #94A24 #Algorithms and Data Compression #Computability, Logic, AI Algorithms #F.2.3 #FOS: Computer and information sciences #Information Theory (cs.IT) #acm:68Q30 #acm:94A15 #acm:94A24 #cs.IT #math.IT #msc:68Q30 #msc:94A15 #msc:94A24 #semigroups and automata theory
paper · pdf · doi:10.48550/arxiv.1911.04268
26 pages
arxiv created 2019/11/11 · openalex publication_date 2019/11/11 · arxiv updated 2019/11/12 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
In a lossless compression system with target lengths, a compressor \cal C maps an integer m and a binary string x to an m-bit code p, and if m is sufficiently large, a decompressor \cal D reconstructs x from p. We call a pair (m,x) achievable for (\cal C,\cal D) if this reconstruction is successful. We introduce the notion of an optimal compressor \cal Copt, by the following universality property: For any compressor-decompressor pair (\cal C, \cal D), there exists a decompressor \cal D' such that if (m,x) is achievable for (\cal C,\cal D), then (m+Δ, x) is achievable for (\cal Copt, \cal D'), where Δ is some small value called the overhead. We show that there exists an optimal compressor that has only polylogarithmic overhead and works in probabilistic polynomial time. Differently said, for any pair (\cal C, \cal D), no matter how slow \cal C is, or even if \cal C is non-computable, \cal Copt is a fixed compressor that in polynomial time produces codes almost as short as those of \cal C. The cost is that the corresponding decompressor is slower. We also show that each such optimal compressor can be used for distributed compression, in which case it can achieve optimal compression rates, as given in the Slepian-Wolf theorem, and even for the Kolmogorov complexity variant of this theorem. Moreover, the overhead is logarithmic in the number of sources, and unlike previous implementations of Slepian-Wolf coding, meaningful compression can still be achieved if the number of sources is much larger than the length of the compressed strings.