vix.ing · top · new · best · stats · spec

Importance Sampling of Word Patterns in DNA and Protein Sequences

2008/11/26 by Hock Peng Chan, Chan, Hock Peng, Nancy R. Zhang +3
Biochemistry, Genetics and Molecular Biology · Mathematics · #Applications (stat.AP) #Computation (stat.CO) #FOS: Biological sciences #FOS: Computer and information sciences #Quantitative Methods (q-bio.QM) #q-bio.QM #stat.AP #stat.CO

paper · pdf · doi:10.48550/arxiv.0811.4447

arxiv created 2008/11/26 · arxiv updated 2009/12/01

Abstract

Monte Carlo methods can provide accurate p-value estimates of word counting test statistics and are easy to implement. They are especially attractive when an asymptotic theory is absent or when either the search sequence or the word pattern is too short for the application of asymptotic formulae. Naive direct Monte Carlo is undesirable for the estimation of small probabilities because the associated rare events of interest are seldom generated. We propose instead efficient importance sampling algorithms that use controlled insertion of the desired word patterns on randomly generated sequences. The implementation is illustrated on word patterns of biological interest: Palindromes and inverted repeats, patterns arising from position specific weight matrices and co-occurrences of pairs of motifs.

Related