2023/08/01 by Carlos E. M. Relvas, Asuka Nakata, Guoan Chen +4 · 2 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · Mathematics · #Advanced Clustering Algorithms Research #Artificial intelligence #Bayesian Methods and Mixture Models #CURE data clustering algorithm #Cluster analysis #Clustering high-dimensional data #Computer science #Correlation clustering #Covariate #Data mining #Gene expression and cancer classification #Machine learning #Mathematics #Set (abstract data type)
paper · open access · doi:10.1142/s0219720023500191
published in Journal of Bioinformatics and Computational Biology 21(04), 2350019 (Imperial College Press)
crossref issued 2023/08/01 · crossref published 2023/08/01 · crossref published-print 2023/08/01 · openalex publication_date 2023/08/01 · crossref created 2023/08/20 · crossref published-online 2023/09/08 · crossref deposited 2023/09/18 · openalex created_date 2025/10/10 · crossref indexed 2026/08/04 · openalex updated_date 2026/08/06
Usually, the clustering process is the first step in several data analyses. Clustering allows identify patterns we did not note before and helps raise new hypotheses. However, one challenge when analyzing empirical data is the presence of covariates, which may mask the obtained clustering structure. For example, suppose we are interested in clustering a set of individuals into controls and cancer patients. A clustering algorithm could group subjects into young and elderly in this case. It may happen because the age at diagnosis is associated with cancer. Thus, we developed CEM-Co, a model-based clustering algorithm that removes/minimizes undesirable covariates' effects during the clustering process. We applied CEM-Co on a gene expression dataset composed of 129 stage I non-small cell lung cancer patients. As a result, we identified a subgroup with a poorer prognosis, while standard clustering algorithms failed.