2013/07/30 by Gilles Celeux, Marie‐Laure Martin‐Magniette, Celeux, Gilles +5
Computer Science · #Advanced Clustering Algorithms Research #Applications (stat.AP) #Bayesian Methods and Mixture Models #Data Mining Algorithms and Applications #FOS: Computer and information sciences
paper · pdf · doi:10.48550/arxiv.1307.7860
openalex publication_date 2013/07/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We compare two major approaches to variable selection in clustering: model\nselection and regularization. Based on previous results, we select the method\nof Maugis et al. (2009b), which modified the method of Raftery and Dean (2006),\nas a current state of the art model selection method. We select the method of\nWitten and Tibshirani (2010) as a current state of the art regularization\nmethod. We compared the methods by simulation in terms of their accuracy in\nboth classification and variable selection. In the first simulation experiment\nall the variables were conditionally independent given cluster membership. We\nfound that variable selection (of either kind) yielded substantial gains in\nclassification accuracy when the clusters were well separated, but few gains\nwhen the clusters were close together. We found that the two variable selection\nmethods had comparable classification accuracy, but that the model selection\napproach had substantially better accuracy in selecting variables. In our\nsecond simulation experiment, there were correlations among the variables given\nthe cluster memberships. We found that the model selection approach was\nsubstantially more accurate in terms of both classification and variable\nselection than the regularization approach, and that both gave more accurate\nclassifications than K-means without variable selection.\n