2014/09/29 by Theodoros Tsiligkaridis, Tsiligkaridis, Theodoros, Keith W. Forsythe +1
Computer Science · Mathematics · #Bayesian Methods and Mixture Models #Data Stream Mining Techniques #FOS: Computer and information sciences #Gaussian Processes and Bayesian Inference #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Methodology (stat.ME) #Statistical Methods and Inference #cs.LG #stat.ME #stat.ML
paper · pdf · doi:10.48550/arxiv.1409.8185
25 pages, To appear in Advances in Neural Information Processing Systems (NIPS) 2015
openalex publication_date 2014/09/29 · arxiv created 2015/09/11 · arxiv updated 2015/09/15 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We develop a sequential low-complexity inference procedure for Dirichlet process mixtures of Gaussians for online clustering and parameter estimation when the number of clusters are unknown a-priori. We present an easily computable, closed form parametric expression for the conditional likelihood, in which hyperparameters are recursively updated as a function of the streaming data assuming conjugate priors. Motivated by large-sample asymptotics, we propose a novel adaptive low-complexity design for the Dirichlet process concentration parameter and show that the number of classes grow at most at a logarithmic rate. We further prove that in the large-sample limit, the conditional likelihood and data predictive distribution become asymptotically Gaussian. We demonstrate through experiments on synthetic and real data sets that our approach is superior to other online state-of-the-art methods.