vix.ing · top · new · best · stats · spec

Imputation Strategies for Clustering Mixed-Type Data with Missing Values

2022/11/26 by Rabea Aschenbruck, Gero Szepannek, Adalbert Wilhelm +1
Computer Science · Mathematics · #Advanced Clustering Algorithms Research #Artificial intelligence #Bayesian Methods and Mixture Models #CURE data clustering algorithm #Categorical variable #Cluster analysis #Computer science #Correlation clustering #Data Management and Algorithms #Data mining #Imputation (statistics) #Mathematics #Missing data #Pooling #Single-linkage clustering #Statistics #Type I and type II errors #k-medians clustering

paper · pdf · doi:10.1007/s00357-022-09422-y

crossref issued 2022/11/26 · crossref published 2022/11/26 · crossref published-online 2022/11/26 · openalex publication_date 2022/11/26 · crossref created 2022/11/26 · crossref published-print 2023/04/01 · crossref deposited 2023/05/15 · openalex created_date 2025/10/10 · crossref indexed 2026/08/05 · openalex updated_date 2026/08/06

Abstract

Abstract Incomplete data sets with different data types are difficult to handle, but regularly to be found in practical clustering tasks. Therefore in this paper, two procedures for clustering mixed-type data with missing values are derived and analyzed in a simulation study with respect to the factors of partition, prototypes, imputed values, and cluster assignment. Both approaches are based on the k-prototypes algorithm (an extension of k-means), which is one of the most common clustering methods for mixed-type data (i.e., numerical and categorical variables). For k-means clustering of incomplete data, the k-POD algorithm recently has been proposed, which imputes the missings with values of the associated cluster center. We derive an adaptation of the latter and additionally present a cluster aggregation strategy after multiple imputation. It turns out that even a simplified and time-saving variant of the presented method can compete with multiple imputation and subsequent pooling.

Citations