vix.ing · top · new · best · stats

Optimizing Prognostic Biomarker Discovery in Pancreatic Cancer Through Hybrid Ensemble Feature Selection and Multi-Omics Data

2025/09/02 by John Zobolas, Zobolas, John, Anne-Marie George +9 · 1 voice
Biochemistry, Genetics and Molecular Biology · Computer Science · Mathematics · #Applications (stat.AP) #FOS: Biological sciences #FOS: Computer and information sciences #Genomics (q-bio.GN) #Machine Learning (cs.LG) #Quantitative Methods (q-bio.QM) #cs.LG #q-bio.GN #q-bio.QM #stat.AP

paper · pdf · doi:10.48550/arxiv.2509.02648

Abstract

Prediction of patient survival using high-dimensional multi-omics data requires systematic feature selection methods that ensure predictive performance, sparsity, and reliability for prognostic biomarker discovery. We developed a hybrid ensemble feature selection (hEFS) approach that combines data subsampling with multiple prognostic models, integrating both embedded and wrapper-based strategies for survival prediction. Omics features are ranked using a voting-theory-inspired aggregation mechanism across models and subsamples, while the optimal number of features is selected via a Pareto front, balancing predictive accuracy and model sparsity without any user-defined thresholds. When applied to multi-omics datasets from three pancreatic cancer cohorts, hEFS identifies significantly fewer and more stable biomarkers compared to the conventional, late-fusion CoxLasso models, while maintaining comparable discrimination performance. Implemented within the open-source mlr3fselect R package, hEFS offers a robust, interpretable, and clinically valuable tool for prognostic modelling and biomarker discovery in high-dimensional survival settings.

Citations

Discussions

Related