vix.ing · top · new · best · stats · spec

Comprehensive benchmarking of large language models for RNA secondary structure prediction

2024/10/21 by Luciano I Zablocki, Leandro A. Bugnon, Zablocki, L. I. +9 · 2 citations
Biochemistry, Genetics and Molecular Biology · #Artificial Intelligence (cs.AI) #Biomolecules (q-bio.BM) #FOS: Biological sciences #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning in Bioinformatics #RNA and protein synthesis mechanisms #RNA modifications and cancer

paper · pdf · doi:10.48550/arxiv.2410.16212

openalex publication_date 2024/10/21 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Inspired by the success of large language models (LLM) for DNA and proteins, several LLM for RNA have been developed recently. RNA-LLM uses large datasets of RNA sequences to learn, in a self-supervised way, how to represent each RNA base with a semantically rich numerical vector. This is done under the hypothesis that obtaining high-quality RNA representations can enhance data-costly downstream tasks. Among them, predicting the secondary structure is a fundamental task for uncovering RNA functional mechanisms. In this work we present a comprehensive experimental analysis of several pre-trained RNA-LLM, comparing them for the RNA secondary structure prediction task in an unified deep learning framework. The RNA-LLM were assessed with increasing generalization difficulty on benchmark datasets. Results showed that two LLM clearly outperform the other models, and revealed significant challenges for generalization in low-homology scenarios.

Cited by

Related