2022/11/26 by Rabea Aschenbruck, Gero Szepannek, Adalbert Wilhelm +1
Computer Science · Mathematics · #Advanced Clustering Algorithms Research #Artificial intelligence #Bayesian Methods and Mixture Models #CURE data clustering algorithm #Categorical variable #Cluster analysis #Computer science #Correlation clustering #Data Management and Algorithms #Data mining #Imputation (statistics) #Mathematics #Missing data #Pooling #Single-linkage clustering #Statistics #Type I and type II errors #k-medians clustering
paper · pdf · doi:10.1007/s00357-022-09422-y
crossref issued 2022/11/26 · crossref published 2022/11/26 · crossref published-online 2022/11/26 · openalex publication_date 2022/11/26 · crossref created 2022/11/26 · crossref published-print 2023/04/01 · crossref deposited 2023/05/15 · openalex created_date 2025/10/10 · crossref indexed 2026/08/05 · openalex updated_date 2026/08/06
Abstract Incomplete data sets with different data types are difficult to handle, but regularly to be found in practical clustering tasks. Therefore in this paper, two procedures for clustering mixed-type data with missing values are derived and analyzed in a simulation study with respect to the factors of partition, prototypes, imputed values, and cluster assignment. Both approaches are based on the k-prototypes algorithm (an extension of k-means), which is one of the most common clustering methods for mixed-type data (i.e., numerical and categorical variables). For k-means clustering of incomplete data, the k-POD algorithm recently has been proposed, which imputes the missings with values of the associated cluster center. We derive an adaptation of the latter and additionally present a cluster aggregation strategy after multiple imputation. It turns out that even a simplified and time-saving variant of the presented method can compete with multiple imputation and subsequent pooling.