2019/12/05 by Tanbin Rahman, Yujia Li, Rahman, Tanbin +7 · 1 citation
Biochemistry, Genetics and Molecular Biology · Computer Science · #Gene expression and cancer classification #Bayesian Methods and Mixture Models #Bioinformatics and Genomic Networks
paper · pdf · doi:10.48550/arxiv.1912.02399
Clustering with variable selection is a challenging yet critical task for\nmodern small-n-large-p data. Existing methods based on sparse Gaussian mixture\nmodels or sparse K-means provide solutions to continuous data. With the\nprevalence of RNA-seq technology and lack of count data modeling for\nclustering, the current practice is to normalize count expression data into\ncontinuous measures and apply existing models with Gaussian assumption. In this\npaper, we develop a negative binomial mixture model with lasso or fused lasso\ngene regularization to cluster samples (small n) with high-dimensional gene\nfeatures (large p). EM algorithm and Bayesian information criterion are used\nfor inference and determining tuning parameters. The method is compared with\nexisting methods using extensive simulations and two real transcriptomic\napplications in rat brain and breast cancer studies. The result shows superior\nperformance of the proposed count data model in clustering accuracy, feature\nselection and biological interpretation in pathways.\n