On the Convergence Properties of the EM Algorithm
1983/03/01 by C. F. Jeff Wu, Changbao Wu · 3,298 citations
Engineering · Mathematics · Physics and Astronomy · #Algorithm #Applied mathematics #Computer science #Control Systems and Identification #Convergence (economics) #Differentiable function #Estimation theory #Expectation–maximization algorithm #Exponential family #Exponential function #Function (biology) #Likelihood function #Limit (mathematics) #Limit of a function #Mathematical analysis #Mathematics #Maximum likelihood #Scientific Research and Discoveries #Sequence (biology) #Statistical and numerical algorithms #Statistics #Weak convergence
paper · pdf · doi:10.1214/aos/1176346060
published in The Annals of Statistics 11(1) (Institute of Mathematical Statistics)
openalex publication_date 1983/03/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/04
Abstract
Two convergence aspects of the EM algorithm are studied: (i) does the EM algorithm find a local maximum or a stationary value of the (incomplete-data) likelihood function? (ii) does the sequence of parameter estimates generated by EM converge? Several convergence results are obtained under conditions that are applicable to many practical situations. Two useful special cases are: (a) if the unobserved complete-data specification can be described by a curved exponential family with compact parameter space, all the limit points of any EM sequence are stationary points of the likelihood function; (b) if the likelihood function is unimodal and a certain differentiability condition is satisfied, then any EM sequence converges to the unique maximum likelihood estimate. A list of key properties of the algorithm is included.
Citations
Cited by
- An Introduction to Bayesian and Frequentist Simulation-Based Inference with Machine Learning
- Convergence of a stochastic approximation version of the EM algorithm
- Fast data inversion for high-dimensional Ornstein-Uhlenbeck processes from noisy measurements
- Information Theory and Statistical Learning
- Divergence-Based Motivation for Online EM and Combining Hidden Variable Models
- A Note on the Convergence of the Gaussian Mean Shift Algorithm
- Probabilistic Constraint Logic Programming. Formal Foundations of Quantitative and Statistical Inference in Constraint-Based Natural Language Processing
- The Dynamic ECME Algorithm
- Statistical and Computational Guarantees for the Baum-Welch Algorithm
- Latent Variable Sequence Identification for Cognitive Models with Neural Network Estimators
- Latent Fisher Discriminant Analysis
- Scalable Bayesian Quantile Regression via EM and INLA: The EM-INLA Algorithm
- Fast expectation-maximization algorithms for spatial generalized linear mixed models
- Composite Difference-Max Programs for Modern Statistical Estimation Problems
- Understanding and Accelerating EM Algorithm's Convergence by Fair Competition Principle and Rate-Verisimilitude Function
- MiCE: Mixture of Contrastive Experts for Unsupervised Image Clustering
- Accelerated Nonparametric Maximum Likelihood Density Deconvolution Using Bernstein Polynomial
- Hopfield Networks is All You Need
- Model-Based Clustering, Discriminant Analysis, and Density Estimation
- Belief Net: A Filter-Based Framework for Learning Hidden Markov Models from Observations
- The Matrix Ridge Approximation: Algorithms and Applications
- CLAX: Fast and Flexible Neural Click Models in JAX
- Phase-type fitting of scale functions for spectrally negative Levy processes
- Recursive maximum likelihood identification of jump Markov nonlinear systems
- Particle-based likelihood inference in partially observed diffusion processes using generalised Poisson estimators
- The expectation-maximization algorithm
- Scaling laws of bacterial and archaeal plasmids
- Singularity, Misspecification, and the Convergence Rate of EM
- Fast Network Community Detection with Profile-Pseudo Likelihood Methods
- Ten Steps of EM Suffice for Mixtures of Two Gaussians
- Fair Marriage Principle and Initialization Map for the EM Algorithm
- EM Converges for a Mixture of Many Linear Regressions
- Efficient preconditioned stochastic gradient descent for estimation in latent variable models
- Sinkhorn EM: An Expectation-Maximization algorithm based on entropic optimal transport
- Statistical Guarantees for Estimating the Centers of a Two-component Gaussian Mixture by EM
- On the Analysis of EM for truncated mixtures of two Gaussians
- A Semiparametric Gaussian Mixture Model with Spatial Dependence and Its Application to Whole-Slide Image Clustering Analysis
- Tobit models: A survey
- Model-based Sparse Coding beyond Gaussian Independent Model
- High-Dimensional Differentially-Private EM Algorithm: Methods and Near-Optimal Statistical Guarantees
- A semiparametric inference to regression analysis with missing covariates in survey data
- Separating the what and how of compositional computation to enable reuse and continual learning
- A Frequentist Statistical Introduction to Variational Inference, Autoencoders, and Diffusion Models
- Bessel regression model: Robustness to analyze bounded data
- EMFlow: Data Imputation in Latent Space via EM and Deep Flow Models
- Joint Channel Estimation and Channel Decoding in Physical-Layer Network Coding Systems: An EM-BP Factor Graph Framework
- Estimating Well-Performing Bayesian Networks using Bernoulli Mixtures
- Benefits of over-parameterization with EM
- Users as Annotators: LLM Preference Learning from Comparison Mode
- Consistency of the MLE under mixture models
- VIREL: A Variational Inference Framework for Reinforcement Learning
- MAGIC: Multi-task Gaussian process for joint imputation and classification in healthcare time series
- Truncated Inference for Latent Variable Optimization Problems: Application to Robust Estimation and Learning
- On the EM-Tau algorithm: a new EM-style algorithm with partial E-steps
- Mixture Models, Robustness, and Sum of Squares Proofs
- A Unified Approach for Learning the Parameters of Sum-Product Networks
- A Tutorial on the Expectation-Maximization Algorithm Including Maximum-Likelihood Estimation and EM Training of Probabilistic Context-Free Grammars
- Hierarchical Mixtures of Experts and the EM Algorithm
- RL's Razor: Why Online Reinforcement Learning Forgets Less
- Differentiable Expectation-Maximisation and Applications to Gaussian Mixture Model Optimal Transport
- Maximum likelihood estimation of the Markov-switching GARCH model
- Causal Intervention for Weakly-Supervised Semantic Segmentation
- Linear-Time Probabilistic Solutions of Boundary Value Problems
- Convergence of the EM Algorithm for Gaussian Mixtures with Unbalanced Mixing Coefficients
- An Imputation-Consistency Algorithm for High-Dimensional Missing Data Problems and Beyond
- Learning dynamical systems with particle stochastic approximation EM
- Statistical Convergence of the EM Algorithm on Gaussian Mixture Models
- Nonparametric learning of stochastic differential equations from sparse and noisy data
- Enabling low-power massive MIMO with ternary ADCs for AIoT sensing
- Gaussian process regression for survival time prediction with genome-wide gene expression
- A Generative Model for Exploring Structure Regularities in Attributed Networks
- Density Estimation from Aggregated Data with Integrated Auxiliary Information: Estimating Population Densities with Geospatial Data
- The ECME Algorithm Using Factor Analysis for DOA Estimation in Nonuniform Noise
- Full Convergence of the Iterative Bayesian Update and Applications to Mechanisms for Privacy Protection
- Uncertainty Quantification for Large-Scale Deep Networks via Post-StoNet Modeling
- Max-Affine Regression: Provable, Tractable, and Near-Optimal Statistical Estimation
- Maximum multinomial likelihood estimation in compound mixture model with application to malaria study
- Alternating Bregman projections and convergence of the EM algorithm
- An em algorithm for quantum Boltzmann machines
- Learning the Value Systems of Societies from Preferences
- Joint Inference on Truth/Rumor and Their Sources in Social Networks
- Truncated Gaussian Noise Estimation in State-Space Models
- Model uncertainty estimation using the expectation maximization algorithm and a particle flow filter
- Comparing EM with GD in Mixture Models of Two Components
- Parallel Clustering of Graphs for Anonymization and Recommender Systems
- Minimax Optimal Convergence Rates for Estimating Ground Truth from Crowdsourced Labels
- Integrating Probabilistic Rules into Neural Networks: A Stochastic EM Learning Algorithm
- Least squares moment identification of binary regression mixtures models
- Settling the Polynomial Learnability of Mixtures of Gaussians
- Escaping Poor Local Minima in Large Scale Robust Estimation
- A mixture model approach to infer land-use influence on point referenced water quality
- Forward-reverse EM algorithm for Markov chains: convergence and numerical analysis
- Observational nonidentifiability, generalized likelihood and free energy
- Fast Incremental Expectation Maximization for finite-sum optimization: nonasymptotic convergence
- Lasso Regression: Estimation and Shrinkage via Limit of Gibbs Sampling
- Loss Tomography from Tree Topologies to General Topologies
- Fitting large mixture models using stochastic component selection
- An adaptive independence sampler MCMC algorithm for infinite dimensional Bayesian inferences
- A generalized EMS algorithm for model selection with incomplete data
- Structured Convolution Matrices for Energy-efficient Deep learning
- Causal Discovery for Irregularly Time Series with Consistency Guarantees
- Joint Phase Tracking and Channel Decoding for OFDM Physical-Layer Network Coding
- Latent Part-of-Speech Sequences for Neural Machine Translation
- Shared Subspace Models for Multi-Group Covariance Estimation
- A Mixture Model for the Regression Analysis of Competing Risks Data
- Space–Time Modelling of Precipitation by Using a Hidden Markov Model and Censored Gaussian Distributions
- Robust Dictionary based Data Representation
- Improved seeding strategies for k-means and k-GMM
- The EM Algorithm gives Sample-Optimality for Learning Mixtures of Well-Separated Gaussians
- Robust Group Anomaly Detection for Quasi-Periodic Network Time Series
- C-mix: a high dimensional mixture model for censored durations, with applications to genetic data
- Non-asymptotic Analysis of Biased Stochastic Approximation Scheme
- False discovery rate control under reduced precision computation for analysis of neuroimaging data
- Analysis of a Generalized Expectation-Maximization Algorithm for\n Gaussian Mixture Models: A Control Systems Perspective
- Assessing Wikipedia-Based Cross-Language Retrieval Models
- The EM Algorithm is Adaptively-Optimal for Unbalanced Symmetric Gaussian Mixtures
- Cumulative Logit Ordinal Regression with Proportional Odds under Nonignorable Missing Response -- Application to Phase III Trial
- Distributed Learning of Finite Gaussian Mixtures
- Contrasting Multiple Social Network Autocorrelations for Binary Outcomes, With Applications To Technology Adoption
- Space Alternating Penalized Kullback Proximal Point Algorithms for Maximizing Likelihood with Nondifferentiable Penalty
- Parameter recovery in two-component contamination mixtures: the \mathbbL2 strategy
- Envelope Methods with Ignorable Missing Data
- Normal/Independent Distributions and Their Applications in Robust Regression
- High Dimensional Expectation-Maximization Algorithm: Statistical Optimization and Asymptotic Normality
- An EM Algorithm for Continuous-time Bivariate Markov Chains
- Modeling and estimating skewed and heavy-tailed populations via unsupervised mixture models
- Broad Spectrum Structure Discovery in Large-Scale Higher-Order Networks
- Modeling Bimodal Discrete Data Using Conway-Maxwell-Poisson Mixture Models
- Learning higher-order sequential structure with cloned HMMs
- Limits of accuracy for parameter estimation and localisation in Single-Molecule Microscopy via sequential Monte Carlo methods
- Two convergence results for an alternation maximization procedure
- A Renewal Model of Intrusion
- Rethinking Location Privacy for Unknown Mobility Behaviors
- Asynchronous Linear Modulation Classification with Multiple Sensors via Generalized EM Algorithm
- From Shannon's Channel to Semantic Channel via New Bayes' Formulas for Machine Learning
- Alternating Minimization for Mixed Linear Regression
- Multivariate Small-Area Estimation for Mixed-type Response Variables with Item Nonresponse
- Empar: EM-based algorithm for parameter estimation of Markov models on trees
- PULasso: High-dimensional variable selection with presence-only data
- A new class of conditional Markov jump processes with regime switching and path dependence: properties and maximum likelihood estimation
- Restoring Calibration for Aligned Large Language Models: A Calibration-Aware Fine-Tuning Approach
- ECME Thresholding Methods for Sparse Signal Reconstruction
- Estimation of parametric and semiparametric mixture models using phi-divergences
- Analysis of Computer Experiments with Functional Response
- Sequential Monte Carlo smoothing with application to parameter estimation in nonlinear state space models
- A Latent Mixture Model for Heterogeneous Causal Mechanisms in Mendelian Randomization
- Nonparametric estimation of a distribution function under biased sampling and censoring
- Markovian Arrival Process Parameter Estimation of Quasi-birth-death Queueing Systems with Utilization Data
- Nonlinear panel data estimation via quantile regressions
- Statistical inference for epidemic processes in a homogeneous community (Part IV of the book Stochastic Epidemic Models and Inference)
- Hypothesis test for normal mixture models: The EM approach
- New consistent and asymptotically normal estimators for random graph mixture models
- The Ames Salmonella/microsome mutagenicity assay
- ℓ1-penalization for mixture regression models
- Generative Modeling of Discrete Latent Structures via Dynamic Policy Gradients
- Using Bayesian Model Averaging to Calibrate Forecast Ensembles
- Probabilistic Quantitative Precipitation Forecasting Using Bayesian Model Averaging
- Probabilistic Wind Vector Forecasting Using Ensembles and Bayesian Model Averaging
- Joint Model and Data Sparsification via the Marginal Likelihood
- Least absolute deviations estimation via the EM algorithm
- Maximum likelihood estimation in nonlinear mixed effects models
- Bayesian Sigmoid-Type Time Series Forecasting with Missing Data for Greenhouse Crops
- Global Likelihood Optimization Via the Cross-Entropy Method, with an Application to Mixture Models
- A new method for the estimation of variance matrix with prescribed zeros in nonlinear mixed effects models
- Alternating Minimization Algorithm with Automatic Relevance Determination for Transmission Tomography under Poisson Noise
- Cascade of phase transitions for multiscale clustering
- Gene and haplotype frequencies for the loci hLA-A, hLA-B, and hLA-DR based on over 13,000 german blood donors
- Practical Learning of Predictive State Representations
- The Bernstein Function: A Unifying Framework of Nonconvex Penalization in Sparse Estimation
- Likelihood non-Gaussianity in large-scale structure analyses
Related