2026/08/05 by Vikram Aithal, Ajit Kumar, Ambika Sharma
Mathematics · #math.ST #math.AT #stat.TH #msc:62R40
20 Pages, 5 figures
arxiv created 2026/08/05 · arxiv updated 2026/08/06
Topological data analysis (TDA) uses topological techniques to extract meaningful shape-based information from complex datasets. Clustering is a central problem in data analysis, and there has been considerable recent interest in understanding how TDA can inform it. Existing approaches either cluster persistence diagrams directly under Wasserstein-type distances, which is computationally expensive, or use vector representations of diagrams. We propose a kernel k-means algorithm built on a convex combination of sliced Wasserstein (SW) kernels, one for each homology under consideration. Unlike other vector representations of persistence diagrams, the SW kernel is both stable and discriminative with respect to the 1-Wasserstein distance. The method outperforms the baselines on two benchmark datasets and remains competitive on a third synthetic dataset, while being computationally efficient. It also outperforms both a single SW kernel on the union of all homology groups and an SW kernel computed directly on the point clouds. The convex combination assigns an interpretable weight to each q-homology kernel. We further validate that the weights identify the discriminating homology.