vix.ing · top · new · best · stats

Selective inference for k-means clustering

2022/03/29 by Yiqun T. Chen, Daniela Witten, Chen, Yiqun T. +2 · 4 citations
Biochemistry, Genetics and Molecular Biology · Mathematics · Medicine · #FOS: Computer and information sciences #Gene expression and cancer classification #Machine Learning (stat.ML) #Methodology (stat.ME) #SARS-CoV-2 detection and testing #Single-cell and spatial transcriptomics #stat.ME #stat.ML

paper · pdf · doi:10.48550/arxiv.2203.15267

arxiv created 2022/03/29 · openalex publication_date 2022/03/29 · arxiv updated 2022/03/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We consider the problem of testing for a difference in means between clusters of observations identified via k-means clustering. In this setting, classical hypothesis tests lead to an inflated Type I error rate. To overcome this problem, we take a selective inference approach. We propose a finite-sample p-value that controls the selective Type I error for a test of the difference in means between a pair of clusters obtained using k-means clustering, and show that it can be efficiently computed. We apply our proposal in simulation, and on hand-written digits data and single-cell RNA-sequencing data.

Citations

Cited by

Related