2025/01/01 by Sebastian Staab, Kim-Isabelle Mayer, Anny Cárdenas +3 · 1 voice
Biochemistry, Genetics and Molecular Biology · Computer Science · #Advanced Clustering Algorithms Research #Bioinformatics and Genomic Networks #Gene expression and cancer classification
paper · doi:10.1093/ismeco/ycaf174
openalex publication_date 2025/01/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29
The rapid advancement of technologies and methods in the life sciences has significantly increased the availability of big data, presenting new challenges for its analysis. Microbiome datasets, in particular, are characterized by extensive feature sets with defined but complex hierarchical structures that are often overlooked or underutilized. Here we introduce a novel metric, UniCor, to identify UNIquely CORrelated eNtities (UNICORNs) in quantitative datasets associated with continuous target variables. These datasets may include microbiome community structures in relation to environmental factors (e.g., temperature, pH, salinity) or biotic variables (e.g., thermal tolerance, oxidative stress). The UniCor metric combines the uniqueness of a given feature within a dataset with its correlation to a target variable of interest. To further enhance its utility, we developed a propagation algorithm (UniCorP), which exploits inherent dataset hierarchies, such as taxonomic levels in microbiome datasets, by selecting and propagating features based on their UniCor metric. Using bacterial community datasets with hierarchical taxonomic annotations and various continuous environmental variables, we demonstrate the ability of the novel metric to reduce features and increase predictive performance in cross-validated Random Forest Regressions. After propagating features with UniCorP and enriching the hierarchical levels with UNICORNs, the predictive performance consistently outperformed control trials for taxonomic aggregation, even at the least granular hierarchical level, allowing a substantial reduction of the feature space. We also compared the metric to existing methods for feature aggregation, showing that it offers stable, competitive predictive performance and feature reduction, within a simple and adaptable framework.