vix.ing · top · new · best · stats · spec

On the Persistence of Clustering Solutions and True Number of Clusters\n in a Dataset

2018/10/31 by Amber Srivastava, Mayank Baranwal, Srivastava, Amber +3
Business, Management and Accounting · Computer Science · Physics and Astronomy · #Advanced Clustering Algorithms Research #Artificial Intelligence (cs.AI) #Complex Network Analysis Techniques #Customer churn and segmentation #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)

paper · pdf · doi:10.48550/arxiv.1811.00102

openalex publication_date 2018/10/31 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Typically clustering algorithms provide clustering solutions with\nprespecified number of clusters. The lack of a priori knowledge on the true\nnumber of underlying clusters in the dataset makes it important to have a\nmetric to compare the clustering solutions with different number of clusters.\nThis article quantifies a notion of persistence of clustering solutions that\nenables comparing solutions with different number of clusters. The persistence\nrelates to the range of data-resolution scales over which a clustering solution\npersists; it is quantified in terms of the maximum over two-norms of all the\nassociated cluster-covariance matrices. Thus we associate a persistence value\nfor each element in a set of clustering solutions with different number of\nclusters. We show that the datasets where natural clusters are a priori known,\nthe clustering solutions that identify the natural clusters are most persistent\n- in this way, this notion can be used to identify solutions with true number\nof clusters. Detailed experiments on a variety of standard and synthetic\ndatasets demonstrate that the proposed persistence-based indicator outperforms\nthe existing approaches, such as, gap-statistic method, X-means, G-means,\nPG-means, dip-means algorithms and information-theoretic method, in\naccurately identifying the clustering solutions with true number of clusters.\nInterestingly, our method can be explained in terms of the phase-transition\nphenomenon in the deterministic annealing algorithm, where the number of\ndistinct cluster centers changes (bifurcates) with respect to an annealing\nparameter.\n

Related