On the Convergence Properties of the EM Algorithm
1983/03/01 by C. F. Jeff Wu, Changbao Wu · 80 citations
Engineering · Mathematics · Physics and Astronomy · #Control Systems and Identification #Statistical and numerical algorithms #Scientific Research and Discoveries
paper · pdf · doi:10.1214/aos/1176346060
Abstract
Two convergence aspects of the EM algorithm are studied: (i) does the EM algorithm find a local maximum or a stationary value of the (incomplete-data) likelihood function? (ii) does the sequence of parameter estimates generated by EM converge? Several convergence results are obtained under conditions that are applicable to many practical situations. Two useful special cases are: (a) if the unobserved complete-data specification can be described by a curved exponential family with compact parameter space, all the limit points of any EM sequence are stationary points of the likelihood function; (b) if the likelihood function is unimodal and a certain differentiability condition is satisfied, then any EM sequence converges to the unique maximum likelihood estimate. A list of key properties of the algorithm is included.
Citations
Cited by
- An Introduction to Bayesian and Frequentist Simulation-Based Inference with Machine Learning
- Convergence of a stochastic approximation version of the EM algorithm
- Fast data inversion for high-dimensional Ornstein-Uhlenbeck processes from noisy measurements
- Information Theory and Statistical Learning
- Divergence-Based Motivation for Online EM and Combining Hidden Variable\n Models
- A Note on the Convergence of the Gaussian Mean Shift Algorithm
- Probabilistic Constraint Logic Programming. Formal Foundations of Quantitative and Statistical Inference in Constraint-Based Natural Language Processing
- The Dynamic ECME Algorithm
- Statistical and Computational Guarantees for the Baum-Welch Algorithm
- Latent Variable Sequence Identification for Cognitive Models with Neural Network Estimators
- Latent Fisher Discriminant Analysis
- Scalable Bayesian Quantile Regression via EM and INLA: The EM-INLA Algorithm
- Fast expectation-maximization algorithms for spatial generalized linear mixed models
- Composite Difference-Max Programs for Modern Statistical Estimation Problems
- Understanding and Accelerating EM Algorithm's Convergence by Fair Competition Principle and Rate-Verisimilitude Function
- MiCE: Mixture of Contrastive Experts for Unsupervised Image Clustering
- Accelerated Nonparametric Maximum Likelihood Density Deconvolution Using Bernstein Polynomial
- Hopfield Networks is All You Need
- Model-Based Clustering, Discriminant Analysis, and Density Estimation
- Belief Net: A Filter-Based Framework for Learning Hidden Markov Models from Observations
- The Matrix Ridge Approximation: Algorithms and Applications
- CLAX: Fast and Flexible Neural Click Models in JAX
- Phase-type fitting of scale functions for spectrally negative Levy processes
- Recursive maximum likelihood identification of jump Markov nonlinear systems
- Particle-based likelihood inference in partially observed diffusion\n processes using generalised Poisson estimators
- The expectation-maximization algorithm
- Scaling laws of bacterial and archaeal plasmids
- Singularity, Misspecification, and the Convergence Rate of EM
- Fast Network Community Detection With Profile-Pseudo Likelihood Methods
- Ten Steps of EM Suffice for Mixtures of Two Gaussians
- Fair Marriage Principle and Initialization Map for the EM Algorithm
- EM Converges for a Mixture of Many Linear Regressions
- Efficient preconditioned stochastic gradient descent for estimation in latent variable models
- Sinkhorn EM: An Expectation-Maximization algorithm based on entropic optimal transport
- Statistical Guarantees for Estimating the Centers of a Two-component Gaussian Mixture by EM
- On the Analysis of EM for truncated mixtures of two Gaussians
- A Semiparametric Gaussian Mixture Model with Spatial Dependence and Its Application to Whole-Slide Image Clustering Analysis
- Tobit models: A survey
- Model-based Sparse Coding beyond Gaussian Independent Model
- High-Dimensional Differentially-Private EM Algorithm: Methods and Near-Optimal Statistical Guarantees
- A semiparametric inference to regression analysis with missing covariates in survey data
- Separating the what and how of compositional computation to enable reuse and continual learning
- A Frequentist Statistical Introduction to Variational Inference, Autoencoders, and Diffusion Models
- Bessel regression model: Robustness to analyze bounded data
- EMFlow: Data Imputation in Latent Space via EM and Deep Flow Models
- Joint Channel Estimation and Channel Decoding in Physical-Layer Network Coding Systems: An EM-BP Factor Graph Framework
- Estimating Well-Performing Bayesian Networks using Bernoulli Mixtures
- Benefits of over-parameterization with EM
- Users as Annotators: LLM Preference Learning from Comparison Mode
- Consistency of the MLE under mixture models
- VIREL: A Variational Inference Framework for Reinforcement Learning
- MAGIC: Multi-task Gaussian process for joint imputation and classification in healthcare time series
- Truncated Inference for Latent Variable Optimization Problems: Application to Robust Estimation and Learning
- On the EM-Tau algorithm: a new EM-style algorithm with partial E-steps
- Mixture Models, Robustness, and Sum of Squares Proofs
- A Unified Approach for Learning the Parameters of Sum-Product Networks
- A Tutorial on the Expectation-Maximization Algorithm Including Maximum-Likelihood Estimation and EM Training of Probabilistic Context-Free Grammars
- Hierarchical Mixtures of Experts and the EM Algorithm
- RL's Razor: Why Online Reinforcement Learning Forgets Less
- Differentiable Expectation-Maximisation and Applications to Gaussian Mixture Model Optimal Transport
- Maximum likelihood estimation of the Markov-switching GARCH model
- Causal Intervention for Weakly-Supervised Semantic Segmentation
- Linear-Time Probabilistic Solutions of Boundary Value Problems
- Convergence of the EM Algorithm for Gaussian Mixtures with Unbalanced\n Mixing Coefficients
- An Imputation-Consistency Algorithm for High-Dimensional Missing Data Problems and Beyond
- Learning dynamical systems with particle stochastic approximation EM
- Statistical Convergence of the EM Algorithm on Gaussian Mixture Models
- Nonparametric learning of stochastic differential equations from sparse and noisy data
- Enabling low-power massive MIMO with ternary ADCs for AIoT sensing
- Gaussian process regression for survival time prediction with genome-wide gene expression
- A Generative Model for Exploring Structure Regularities in Attributed Networks
- Density Estimation from Aggregated Data with Integrated Auxiliary Information: Estimating Population Densities with Geospatial Data
- The ECME Algorithm Using Factor Analysis for DOA Estimation in Nonuniform Noise
- Full Convergence of the Iterative Bayesian Update and Applications to Mechanisms for Privacy Protection
- Uncertainty Quantification for Large-Scale Deep Networks via Post-StoNet Modeling
- Max-Affine Regression: Provable, Tractable, and Near-Optimal Statistical\n Estimation
- Alternating Bregman projections and convergence of the EM algorithm
- An em algorithm for quantum Boltzmann machines
- Learning the Value Systems of Societies from Preferences
- Joint Inference on Truth/Rumor and Their Sources in Social Networks
Related