vix.ing · top · new · best · stats · spec

Extended feature allocation models

2025/02/14 by Mario Beraha, Beraha, Mario, Federico Camerlenghi +3 · 1 citation
Computer Science · #Data Management and Algorithms #FOS: Computer and information sciences #FOS: Mathematics #Image Retrieval and Classification Techniques #Machine Learning and Data Classification #Methodology (stat.ME) #Statistics Theory (math.ST)

paper · pdf · doi:10.48550/arxiv.2502.10257

openalex publication_date 2025/02/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/01

Abstract

Feature allocation models are Bayesian nonparametric tools tailored to data in which each observation can simultaneously exhibit multiple characteristics, or features. A fundamental limitation of standard formulations is that feature labels are assumed to be independent and identically distributed, and therefore play no role in posterior inference. The present paper introduces a unified Bayesian framework for extended feature allocation models, in which feature labels and proportions are modeled jointly, thereby enabling the simultaneous discovery of features and learning of dependencies among their labels. Building on point process theory, we develop a full Bayesian analysis of these models. Within this general setting, we also characterize previously proposed priors as those leading to poor predictive distributions, which cannot capture label dependencies and are insensitive to the observed frequency spectrum. Our methodology is designed to move beyond such standard formulations by leveraging the information carried by feature labels. We demonstrate the usefulness of our approach by introducing: (i) a Cox process prior that clusters genomic variant embeddings while predicting new variants and new variant clusters; (ii) a determinantal point process prior for repeated forest surveys, where prediction concerns both the number and the locations of unobserved trees.

Cited by

Related