vix.ing · top · new · best · stats · spec

Principal component analysis based clustering for high-dimension, low-sample-size data

2015/03/16 by Kazuyoshi Yata, Makoto Aoshima, Yata, Kazuyoshi +1
Computer Science · Mathematics · #62H25 #62H30 #Bayesian Methods and Mixture Models #FOS: Mathematics #Face and Expression Recognition #Random Matrices and Applications #Statistics Theory (math.ST) #math.ST #msc:62H25 #msc:62H30 #stat.TH

paper · pdf · doi:10.48550/arxiv.1503.04525

19 pages

arxiv created 2015/03/16 · openalex publication_date 2015/03/16 · arxiv updated 2015/03/17 · openalex created_date 2016/06/24 · openalex updated_date 2026/07/28

Abstract

In this paper, we consider clustering based on principal component analysis (PCA) for high-dimension, low-sample-size (HDLSS) data. We give theoretical reasons why PCA is effective for clustering HDLSS data. First, we derive a geometric representation of HDLSS data taken from a two-class mixture model. With the help of the geometric representation, we give geometric consistency properties of sample principal component scores in the HDLSS context. We develop ideas of the geometric representation and geometric consistency properties to multiclass mixture models. We show that PCA can classify HDLSS data under certain conditions in a surprisingly explicit way. Finally, we demonstrate the performance of the clustering by using microarray data sets.

Citations

Related