2011/04/27 by I‐Ping Tu, I-Ping Tu, Yuan-Fu Huang +4
Biochemistry, Genetics and Molecular Biology · Mathematics · Medicine · #Applications (stat.AP) #FOS: Computer and information sciences #Gene expression and cancer classification #Genetic factors in colorectal cancer #Molecular Biology Techniques and Applications #stat.AP
paper · pdf · doi:10.48550/arxiv.1104.5064
16 pages, 1 figure
arxiv created 2011/04/27 · openalex publication_date 2011/04/27 · arxiv updated 2011/04/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
A DNA palindrome is a segment of double-stranded DNA sequence with inver- sion symmetry which may form secondary structures conferring significant biolog- ical functions ranging from RNA transcription to DNA replication. To test if the clusters of DNA palindromes distribute randomly is an interesting bioinformatic problem, where the occurrence rate of the DNA palindromes is a key estimator for setting up a test. The most commonly used statistics for estimating the occur- rence rate for scan statistics is the average rate. However, in our simulation, the average rate may double the null occurrence rate of DNA palindromes due to hot spot regions of 3000 bp's in a herpes virus genome. Here, we propose a formula to estimate the occurrence rate through an analytic derivation under a Markov assumption on DNA sequence. Our simulation study shows that the performance of this method has improved the accuracy and robustness against hot spots, as compared to the commonly used average rate. In addition, we derived analytical formula for the moment-generating functions of various statistics under a Markov model, enabling further calculations of p-values.