2024/12/20 by Erin Franklin, Élise Billoir, Philippe Veber +3 · 1 voice
Biochemistry, Genetics and Molecular Biology · #Gene expression and cancer classification #Machine Learning in Bioinformatics
paper · doi:10.1101/2024.12.18.627334
openalex publication_date 2024/12/20 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/22
ABSTRACT Interpreting transcriptomic data presents significant challenges, particularly in non-targeted approaches. While modern functional enrichment methods are well-suited for experimental designs involving two conditions, they are less applicable to data series. In this context, we developed Cluefish, a free and open-source, semi-automated R workflow designed for untargeted, comprehensive biological interpretation of transcriptomic data series. Cluefish applies over-representation analysis on pre-clustered protein-protein interaction networks, using clusters as anchors to identify smaller, more specific biological functions. Innovative features, including cluster merging and recovery of isolated genes through shared biological contexts, enable a more complete exploration of the data. We applied Cluefish to an in-house dataset with zebrafish exposed to a dose-gradient of dibutyl phthalate, and to two published toxicology datasets featuring different organisms. Combined with DRomics, a tool for dose-response analysis—Cluefish identified gene clusters deregulated at low doses and linked to biological functions overlooked by the standard approach. Notably, it revealed that retinoid signalling disruption may be the most sensitive pathway affected by dibutyl phthalate during zebrafish development, potentially leading to morphological changes. The Cluefish workflow aims to provide valuable clues for biological hypothesis generation and experimental validation. It is freely available at https://github.com/ellfran-7/cluefish . GRAPHICAL ABSTRACT