vix.ing · top · new · best · stats · spec

Artificial Intelligence in Bulk RNA-Seq: Challenges and Potential Solutions

2026/03/01 by Mostafa Rezapour, Stephanie V. Trefry, Lorreta Aboagyewa Opoku +1 · 1 voice
Biochemistry, Genetics and Molecular Biology · #Cancer-related molecular mechanisms research #Genomics and Phylogenetic Studies #Single-cell and spatial transcriptomics

paper · doi:10.34133/csbj.0039

openalex publication_date 2026/03/01 · openalex created_date 2026/03/17 · openalex updated_date 2026/08/01

Abstract

Bulk RNA sequencing (RNA-seq) produces high-dimensional gene expression data where the number of measured features greatly exceeds the number of available samples, which challenges artificial intelligence (AI)-based modeling. In this setting, models are highly susceptible to overfitting and may fail to generalize across independent datasets when feature dimensionality is not adequately controlled. Feature (gene) selection is therefore essential for reliable inference, particularly when it is performed strictly within training data to prevent information leakage. This review examines how high dimensionality and limited sample size constrain AI-based analysis of bulk RNA-seq data and surveys feature selection strategies used to address these challenges. Emphasis is placed on statistically guided, training-only frameworks, with a focus on generalized linear models with quasi-likelihood F tests and magnitude–altitude scoring (GLMQL-MAS). Across published viral infection studies, GLMQL-MAS yields compact, interpretable gene sets that support robustness, reproducibility, and cross-dataset generalization in bulk transcriptomic modeling.

Citations

Discussions

Related