2024/01/08 by Jiajia Liu, Liu, Jiajia, Mengyuan Yang +12 · 1 voice · 3 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · #Artificial intelligence #Artificial neural network #Computer science #Data science #Focus (optics) #Foundation (evidence) #Language model #Lexical analysis #Machine Learning in Bioinformatics #Machine learning #Natural language #Natural language processing #RNA modifications and cancer #Single-cell and spatial transcriptomics #Topic Modeling #Transformer
paper · pdf · open access · doi:10.1093/bib/bbag367
published in Briefings in Bioinformatics 27(4) (Oxford University Press)
openalex publication_date 2026/07/01 · openalex created_date 2026/07/11 · openalex updated_date 2026/07/23
Large language models (LLMs) are deep learning-based artificial intelligence models that have achieved remarkable success in natural language processing. Typically composed of neural networks with billions of parameters, they are trained on massive unlabeled datasets using self-supervised or semi-supervised learning. Beyond language, LLMs hold immense potential for addressing complex bioinformatics challenges. This review provides a comprehensive overview of transformer-based model applications in genomics, transcriptomics, proteomics, drug discovery, and single-cell analysis. We discuss critical components, including tokenization strategies for diverse biological data, transformer architectures, attention mechanisms, and pretraining approaches. We also survey currently available foundation models and their downstream applications across bioinformatics domains. Finally, we highlight major challenges that remain insufficiently addressed in prior reviews and outline future perspectives and design principles for next-generation biological language models, offering practical guidance for both users and developers.