vix.ing · top · new · best · stats · spec

Unsupervised learning of regression mixture models with unknown number\n of components

2014/09/24 by Faïcel Chamroukhi, Chamroukhi, Faicel
Computer Science · #Advanced Clustering Algorithms Research #Bayesian Methods and Mixture Models #FOS: Computer and information sciences #Face and Expression Recognition #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Methodology (stat.ME)

paper · pdf · doi:10.48550/arxiv.1409.6981

openalex publication_date 2014/09/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Regression mixture models are widely studied in statistics, machine learning\nand data analysis. Fitting regression mixtures is challenging and is usually\nperformed by maximum likelihood by using the expectation-maximization (EM)\nalgorithm. However, it is well-known that the initialization is crucial for EM.\nIf the initialization is inappropriately performed, the EM algorithm may lead\nto unsatisfactory results. The EM algorithm also requires the number of\nclusters to be given a priori; the problem of selecting the number of mixture\ncomponents requires using model selection criteria to choose one from a set of\npre-estimated candidate models. We propose a new fully unsupervised algorithm\nto learn regression mixture models with unknown number of components. The\ndeveloped unsupervised learning approach consists in a penalized maximum\nlikelihood estimation carried out by a robust expectation-maximization (EM)\nalgorithm for fitting polynomial, spline and B-spline regressions mixtures. The\nproposed learning approach is fully unsupervised: 1) it simultaneously infers\nthe model parameters and the optimal number of the regression mixture\ncomponents from the data as the learning proceeds, rather than in a two-fold\nscheme as in standard model-based clustering using afterward model selection\ncriteria, and 2) it does not require accurate initialization unlike the\nstandard EM for regression mixtures. The developed approach is applied to curve\nclustering problems. Numerical experiments on simulated data show that the\nproposed robust EM algorithm performs well and provides accurate results in\nterms of robustness with regard initialization and retrieving the optimal\npartition with the actual number of clusters. An application to real data in\nthe framework of functional data clustering, confirms the benefit of the\nproposed approach for practical applications.\n

Citations

Related