vix.ing · top · new · best · stats · spec

Constructions for Clumps Statistics

2008/04/23 by Frederique Bassino, Frédérique Bassino, Julien Clément +9
Computer Science · Mathematics · #Advanced Combinatorial Mathematics #Algorithms and Data Compression #Discrete Mathematics (cs.DM) #FOS: Computer and information sciences #Information Retrieval (cs.IR) #cs.DM #cs.IR #semigroups and automata theory

paper · pdf · doi:10.48550/arxiv.0804.3671

12 p., 2 figs

arxiv created 2008/04/23 · openalex publication_date 2008/04/23 · arxiv updated 2009/12/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We consider a component of the word statistics known as clump; starting from a finite set of words, clumps are maximal overlapping sets of these occurrences. This parameter has first been studied by Schbath with the aim of counting the number of occurrences of words in random texts. Later work with similar probabilistic approach used the Chen-Stein approximation for a compound Poisson distribution, where the number of clumps follows a law close to Poisson. Presently there is no combinatorial counterpart to this approach, and we fill the gap here. We emphasize the fact that, in contrast with the probabilistic approach which only provides asymptotic results, the combinatorial approach provides exact results that are useful when considering short sequences.

Citations

Related