2016/01/19 by Saverio Ranciati, Cinzia Viroli, Ranciati, Saverio +3
Biochemistry, Genetics and Molecular Biology · Computer Science · #Applications (stat.AP) #Bayesian Methods and Mixture Models #Bioinformatics and Genomic Networks #FOS: Computer and information sciences #Gene expression and cancer classification
paper · pdf · doi:10.48550/arxiv.1601.04879
openalex publication_date 2016/01/19 · openalex created_date 2022/10/01 · openalex updated_date 2026/07/28
Model-based clustering is a technique widely used to group a collection of\nunits into mutually exclusive groups. There are, however, situations in which\nan observation could in principle belong to more than one cluster. In the\ncontext of Next-Generation Sequencing (NGS) experiments, for example, the\nsignal observed in the data might be produced by two (or more) different\nbiological processes operating together and a gene could participate in both\n(or all) of them. We propose a novel approach to cluster NGS discrete data,\ncoming from a ChIP-Seq experiment, with a mixture model, allowing each unit to\nbelong potentially to more than one group: these multiple allocation clusters\ncan be flexibly defined via a function combining the features of the original\ngroups without introducing new parameters. The formulation naturally gives rise\nto a `zero-inflation group' in which values close to zero can be allocated,\nacting as a correction for the abundance of zeros that manifest in this type of\ndata. We take into account the spatial dependency between observations, which\nis described through a latent Conditional Auto-Regressive process that can\nreflect different dependency patterns. We assess the performance of our model\nwithin a simulation environment and then we apply it to ChIP-seq real data.\n