2011/11/25 by Vincent Y. F. Tan, Tan, Vincent Y. F., Cédric Févotte +1 · 3 citations
Computer Science · Engineering · Mathematics · #Algorithm #Artificial intelligence #Artificial neural network #Bayesian probability #Blind Source Separation Techniques #Computer science #Dimensionality reduction #FOS: Computer and information sciences #Face and Expression Recognition #Kullback–Leibler divergence #Machine Learning (stat.ML) #Mathematics #Matrix decomposition #Methodology (stat.ME) #Non-negative matrix factorization #Overfitting #Pattern recognition (psychology) #Robustness (evolution) #Singular value decomposition #Sparse and Compressive Sensing Techniques #stat.ME #stat.ML
paper · pdf · doi:10.48550/arxiv.1111.6085
published in arXiv (Cornell University) (Cornell University) · Accepted by the IEEE Transactions on Pattern Analysis and Machine Intelligence
openalex publication_date 2011/11/25 · arxiv created 2012/10/05 · arxiv updated 2012/10/08 · openalex created_date 2022/10/01 · openalex updated_date 2026/07/28
This paper addresses the estimation of the latent dimensionality in nonnegative matrix factorization (NMF) with the β-divergence. The β-divergence is a family of cost functions that includes the squared Euclidean distance, Kullback-Leibler and Itakura-Saito divergences as special cases. Learning the model order is important as it is necessary to strike the right balance between data fidelity and overfitting. We propose a Bayesian model based on automatic relevance determination in which the columns of the dictionary matrix and the rows of the activation matrix are tied together through a common scale parameter in their prior. A family of majorization-minimization algorithms is proposed for maximum a posteriori (MAP) estimation. A subset of scale parameters is driven to a small lower bound in the course of inference, with the effect of pruning the corresponding spurious components. We demonstrate the efficacy and robustness of our algorithms by performing extensive experiments on synthetic data, the swimmer dataset, a music decomposition example and a stock price prediction task.