2014/02/09 by Niklas Larsson, N. Jesper Larsson, Larsson, N. Jesper
Computer Science · Mathematics · #Advanced Data Compression Techniques #Algorithms and Data Compression #Data Structures and Algorithms (cs.DS) #FOS: Computer and information sciences #Information Theory (cs.IT) #cs.DS #cs.IT #math.IT #semigroups and automata theory
paper · pdf · doi:10.48550/arxiv.1402.1936
arxiv created 2014/02/09 · openalex publication_date 2014/02/09 · arxiv updated 2014/02/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Compression of integer sets and sequences has been extensively studied for settings where elements follow a uniform probability distribution. In addition, methods exist that exploit clustering of elements in order to achieve higher compression performance. In this work, we address the case where enumeration of elements may be arbitrary or random, but where statistics is kept in order to estimate probabilities of elements. We present a recursive subset-size encoding method that is able to benefit from statistics, explore the effects of permuting the enumeration order based on element probabilities, and discuss general properties and possibilities for this class of compression problem.