vix.ing · top · new · best · stats · spec

Poisson approximation for search of rare words in DNA sequences

2007/11/15 by Nicolas Vergne, Vergne, Nicolas, Miguel Abadi +1
Biochemistry, Genetics and Molecular Biology · Mathematics · #60-04 (Secondary) #60F05 #60G10 #92D20 (Primary) #Applications (stat.AP) #DNA and Biological Computing #FOS: Computer and information sciences #FOS: Mathematics #Genomics and Phylogenetic Studies #Probability (math.PR) #RNA and protein synthesis mechanisms #Statistics Theory (math.ST) #math.PR #math.ST #msc:60-04 #msc:60F05 #msc:60G10 #msc:92D20 #stat.AP #stat.TH

paper · pdf · doi:10.48550/arxiv.0711.2382

29 pages, 0 figures

arxiv created 2007/11/15 · openalex publication_date 2007/11/15 · arxiv updated 2009/12/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Using recent results on the occurrence times of a string of symbols in a stochastic process with mixing properties, we present a new method for the search of rare words in biological sequences generally modelled by a Markov chain. We obtain a bound on the error between the distribution of the number of occurrences of a word in a sequence (under a Markov model) and its Poisson approximation. A global bound is already given by a Chen-Stein method. Our approach, the psi-mixing method, gives local bounds. Since we only need the error in the tails of distribution, the global uniform bound of Chen-Stein is too large and it is a better way to consider local bounds. We search for two thresholds on the number of occurrences from which we can regard the studied word as an over-represented or an under-represented one. A biological role is suggested for these over- or under-represented words. Our method gives such thresholds for a panel of words much broader than the Chen-Stein method. Comparing the methods, we observe a better accuracy for the psi-mixing method for the bound of the tails of distribution. We also present the software PANOW (available at http://stat.genopole.cnrs.fr/software/panowdir/) dedicated to the computation of the error term and the thresholds for a studied word.

Related