2021/06/25 by Heejung Shim, Shim, Heejung, Zhengrong Xing +10
Biochemistry, Genetics and Molecular Biology · Mathematics · #FOS: Biological sciences #FOS: Computer and information sciences #Gene expression and cancer classification #Methodology (stat.ME) #Molecular Biology Techniques and Applications #Quantitative Methods (q-bio.QM) #Single-cell and spatial transcriptomics #q-bio.QM #stat.ME
paper · pdf · doi:10.48550/arxiv.2106.13634
24 pages, 4 figures
arxiv created 2021/06/25 · openalex publication_date 2021/06/25 · arxiv updated 2021/06/28 · openalex created_date 2021/07/05 · openalex updated_date 2026/07/28
Estimating and testing for differences in molecular phenotypes (e.g. gene expression, chromatin accessibility, transcription factor binding) across conditions is an important part of understanding the molecular basis of gene regulation. These phenotypes are commonly measured using high-throughput sequencing assays (e.g., RNA-seq, ATAC-seq, ChIP-seq), which provide high-resolution count data that reflect how the phenotypes vary along the genome. Multiple methods have been proposed to help exploit these high-resolution measurements for differential expression analysis. However, they ignore the count nature of the data, instead using normal approximations that work well only for data with large sample sizes or high counts. Here we develop count-based methods to address this problem. We model the data for each sample using an inhomogeneous Poisson process with spatially structured underlying intensity function, and then, building on multi-scale models for the Poisson process, estimate and test for differences in the underlying intensity function across samples (or groups of samples). Using both simulation and real ATAC-seq data we show that our method outperforms previous normal-based methods, especially in situations with small sample sizes or low counts.