vix.ing · top · new · best · stats · spec

Inference on High-Dimensional Sparse Count Data

2015/10/14 by Jyotishka Datta, David B. Dunson, Datta, Jyotishka +1
Biochemistry, Genetics and Molecular Biology · Computer Science · Mathematics · #62C10 #62F15 #Bayesian Methods and Mixture Models #FOS: Computer and information sciences #Genetic Associations and Epidemiology #Methodology (stat.ME) #Statistical Methods and Inference

paper · pdf · doi:10.48550/arxiv.1510.04320

openalex publication_date 2015/10/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

In a variety of application areas, there is a growing interest in analyzing high dimensional sparse count data, with sparsity exhibited by an over-abundance of zeros and small non-zero counts. Existing approaches for analyzing multivariate count data via Poisson or negative binomial log-linear hierarchical models with zero-inflation cannot flexibly adapt to the level and nature of sparsity in the data. We develop a new class of continuous local-global shrinkage priors tailored for sparse counts. Theoretical properties are assessed, including posterior concentration, stronger control on false discoveries in multiple testing, robustness in posterior mean and super-efficiency in estimating the sampling density. Simulation studies illustrate excellent small sample properties relative to competitors. We apply the method to detect rare mutational hotspots in exome sequencing data and to identify cities most impacted by terrorism.

Citations

Related