vix.ing · top · new · best · stats

Leveraging genomic deep learning models for the prediction of non-coding variant effects

2024/11/17 by Pooja Kathail, Ayesha Bajwa, Kathail, Pooja +3 · 1 voice · 5 citations
Biochemistry, Genetics and Molecular Biology · Mathematics · #Artificial intelligence #Biology #Coding (social sciences) #Computational biology #Computer science #Deep learning #Genetics, Bioinformatics, and Biomedical Research #Machine learning #Mathematics #Statistics

paper · pdf · doi:10.48550/arxiv.2411.11158

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2024/11/17 · openalex created_date 2024/11/21 · openalex updated_date 2026/07/28

Abstract

Characterizing non-coding variant function remains an important challenge in human genetics. Genomic deep learning models have emerged as a promising approach to enable in silico prediction of variant effects. These include supervised sequence-to-activity models, which predict molecular phenotypes such as genome-wide chromatin states or gene expression levels directly from DNA sequence, and self-supervised genomic language models. Here, we review progress in leveraging these models for non-coding variant effect prediction. We describe practical considerations for making such predictions and categorize the types of ground truth data used to evaluate variant effect predictions, providing insight into the settings in which current models are most useful. Our Review highlights key considerations for practitioners and opportunities for improvement in model development and evaluation.

Cited by

Discussions

Related