vix.ing · top · new · best · stats · spec

Efficient EM Training of Gaussian Mixtures with Missing Data

2012/09/04 by Olivier Delalleau, Aaron Courville, Delalleau, Olivier +3
Computer Science · Mathematics · #Algorithms and Data Compression #Bayesian Methods and Mixture Models #Data Mining Algorithms and Applications #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #cs.LG #stat.ML

paper · pdf · doi:10.48550/arxiv.1209.0521

openalex publication_date 2012/09/04 · arxiv created 2018/01/08 · arxiv updated 2018/01/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

In data-mining applications, we are frequently faced with a large fraction of missing entries in the data matrix, which is problematic for most discriminant machine learning algorithms. A solution that we explore in this paper is the use of a generative model (a mixture of Gaussians) to compute the conditional expectation of the missing variables given the observed variables. Since training a Gaussian mixture with many different patterns of missing values can be computationally very expensive, we introduce a spanning-tree based algorithm that significantly speeds up training in these conditions. We also observe that good results can be obtained by using the generative model to fill-in the missing values for a separate discriminant learning algorithm.

Citations

Related