2014/08/14 by Creighton Heaukulani, Heaukulani, Creighton, David A. Knowles +3
Biochemistry, Genetics and Molecular Biology · Computer Science · Mathematics · #Bayesian Methods and Mixture Models #Bioinformatics and Genomic Networks #FOS: Computer and information sciences #Gene expression and cancer classification #Machine Learning (stat.ML) #stat.ML
paper · pdf · doi:10.48550/arxiv.1408.3378
43 pages, 13 figures. Major revision to the proof of Thm. 2. Large portions of Chs. 2 & 4 moved into the appendix. Added Fig. 4. Revisions throughout
openalex publication_date 2014/08/14 · arxiv created 2015/04/03 · arxiv updated 2015/04/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We define the beta diffusion tree, a random tree structure with a set of leaves that defines a collection of overlapping subsets of objects, known as a feature allocation. A generative process for the tree structure is defined in terms of particles (representing the objects) diffusing in some continuous space, analogously to the Dirichlet diffusion tree (Neal, 2003), which defines a tree structure over partitions (i.e., non-overlapping subsets) of the objects. Unlike in the Dirichlet diffusion tree, multiple copies of a particle may exist and diffuse along multiple branches in the beta diffusion tree, and an object may therefore belong to multiple subsets of particles. We demonstrate how to build a hierarchically-clustered factor analysis model with the beta diffusion tree and how to perform inference over the random tree structures with a Markov chain Monte Carlo algorithm. We conclude with several numerical experiments on missing data problems with data sets of gene expression microarrays, international development statistics, and intranational socioeconomic measurements.