2021/08/17 by Yi Li, Yan Song, Li, Yi +3
Computer Science · Decision Sciences · #Advanced Clustering Algorithms Research #Data Management and Algorithms #Data Quality and Management #FOS: Computer and information sciences #Machine Learning (cs.LG)
paper · pdf · doi:10.48550/arxiv.2108.07383
openalex publication_date 2021/08/17 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We study the problem of learning to cluster data points using an oracle which can answer same-cluster queries. Different from previous approaches, we do not assume that the total number of clusters is known at the beginning and do not require that the true clusters are consistent with a predefined objective function such as the K-means. These relaxations are critical from the practical perspective and, meanwhile, make the problem more challenging. We propose two algorithms with provable theoretical guarantees and verify their effectiveness via an extensive set of experiments on both synthetic and real-world data.