Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing
2021/10/15 by 裕二 池谷, Robert Tinn, Hao Cheng +6 · 94 citations
Computer Science · Biochemistry, Genetics and Molecular Biology · #Topic Modeling #Natural Language Processing Techniques #Biomedical Text Mining and Ontologies
paper · doi:10.1145/3458754
openalex created_date 2020/08/07 · openalex publication_date 2021/10/15 · openalex updated_date 2026/07/29
Abstract
Pretraining large neural language models, such as BERT, has led to impressive gains on many natural language processing (NLP) tasks. However, most pretraining efforts focus on general domain corpora, such as newswire and Web. A prevailing assumption is that even domain-specific pretraining can benefit by starting from general-domain language models. In this article, we challenge this assumption by showing that for domains with abundant unlabeled text, such as biomedicine, pretraining language models from scratch results in substantial gains over continual pretraining of general-domain language models. To facilitate this investigation, we compile a comprehensive biomedical NLP benchmark from publicly available datasets. Our experiments show that domain-specific pretraining serves as a solid foundation for a wide range of biomedical NLP tasks, leading to new state-of-the-art results across the board. Further, in conducting a thorough evaluation of modeling choices, both for pretraining and task-specific fine-tuning, we discover that some common practices are unnecessary with BERT models, such as using complex tagging schemes in named entity recognition. To help accelerate research in biomedical NLP, we have released our state-of-the-art pretrained and task-specific models for the community, and created a leaderboard featuring our BLURB benchmark (short for Biomedical Language Understanding & Reasoning Benchmark) at https://aka.ms/BLURB .
Cited by
- Artificial intelligence in bioinformatics: a survey
- Machine learning based screening of potential paper mill publications in cancer research: methodological and cross sectional study
- Improving large language models for clinical named entity recognition via prompt engineering
- Large language models for biomedicine: foundations, opportunities, challenges, and best practices
- BioHiCL: Hierarchical Multi-Label Contrastive Learning for Biomedical Retrieval with MeSH Labels
- Large language models encode clinical knowledge
- Generative AI for Healthcare: Fundamentals, Challenges, and Perspectives
- Revealing the Paper Mill Iceberg: AI-Based Screening of Cancer Research Publications
- Accurate and Scalable Multimodal Pathology Retrieval via Attentive Vision-Language Alignment
- Comparison of biomedical relationship extraction methods and models for knowledge graph creation
- Large language models in medicine
- ECG-LLM -- training and evaluation of domain-specific large language models for electrocardiography
- ACTG-ARL: Differentially Private Conditional Text Generation with RL-Boosted Control
- BRAINCELL-AID: An Agentic AI Created Brain Cell Type Resource for Community Annotation
- OG-Rank: Learning to Rank Fast and Slow with Uncertainty and Reward-Trend Guided Adaptive Exploration
- MOSAIC: Masked Objective with Selective Adaptation for In-domain Contrastive Learning
- Open WebUI: An Open, Extensible, and Usable Interface for AI Interaction
- Generalist vs Specialist Time Series Foundation Models: Investigating Potential Emergent Behaviors in Assessing Human Health Using PPG Signals
- JEDA: Query-Free Clinical Order Search from Ambient Dialogues
- Towards Human-Centric Intelligent Treatment Planning for Radiation Therapy
- ProtoTopic: Prototypical Network for Few-Shot Medical Topic Modeling
- MoRA: On-the-fly Molecule-aware Low-Rank Adaptation Framework for LLM-based Multi-Modal Molecular Assistant
- VeritasFi: An Adaptable, Multi-tiered RAG Framework for Multi-modal Financial Question Answering
- Variational Open-Domain Question Answering
- You only need 4 extra tokens: Synergistic Test-time Adaptation for LLMs
- Lightweight Baselines for Medical Abstract Classification: DistilBERT with Cross-Entropy as a Strong Default
- Serialized EHR make for good text representations
- SMedBERT: A Knowledge-Enhanced Pre-trained Language Model with Structured Semantics for Medical Text Mining
- HySim-LLM: Embedding-Weighted Fine-Tuning Bounds and Manifold Denoising for Domain-Adapted LLMs
- Recent Advances in Automated Question Answering In Biomedical Domain
- Annotating publicly-available samples and studies using interpretable modeling of unstructured metadata
- Hierarchical attention graph learning with LLM enhancement for molecular solubility prediction
- BigBIO: A Framework for Data-Centric Biomedical Natural Language Processing
- ModernBERT + ColBERT: Enhancing biomedical RAG through an advanced re-ranking retriever
- Named Entity Recognition in COVID-19 tweets with Entity Knowledge Augmentation
- Memory-Augmented Log Analysis with Phi-4-mini: Enhancing Threat Detection in Structured Security Logs
- Thin Bridges for Drug Text Alignment: Lightweight Contrastive Learning for Target Specific Drug Retrieval
- CliniBench: A Clinical Outcome Prediction Benchmark for Generative and Encoder-Based Language Models
- Decipher-MR: A Vision-Language Foundation Model for 3D MRI Representations
- RAD: Towards Trustworthy Retrieval-Augmented Multi-modal Clinical Diagnosis
- Revisiting Performance Claims for Chest X-Ray Models Using Clinical Context
- The MADRS Pipeline: Supporting Depression Assessment in Clinical Trials
- A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports
- Automated extraction of fungal trophic modes from literature using BioBERT: an open pilot workflow
- MedLLM: An Open Medical Language Model at the Sub-Billion Scale
- Pretrained Transformers for Text Ranking: BERT and Beyond
- Advances in Large Language Models for Medicine
- Are Smaller Open-Weight LLMs Closing the Gap to Proprietary Models for Biomedical Question Answering?
- A Novel Metric for Detecting Memorization in Generative Models for Brain MRI Synthesis
- GRIL: Knowledge Graph Retrieval-Integrated Learning with Large Language Models
- Uncertainty Quantification of Large Language Models using Approximate Bayesian Computation
- Scientific Language Models for Biomedical Knowledge Base Completion: An Empirical Study
- Efficient and Versatile Model for Multilingual Information Retrieval of Islamic Text: Development and Deployment in Real-World Scenarios
- Combining Evidence and Reasoning for Biomedical Fact-Checking
- On the Universality of Deep Contextual Language Models
- Multi Anatomy X-Ray Foundation Model
- Biomedical Hypothesis Explainability with Graph-Based Context Retrieval
- RanAT4BIE: Random Adversarial Training for Biomedical Information Extraction
- GAPrune: Gradient-Alignment Pruning for Domain-Aware Embeddings
- Automated MCQA Benchmarking at Scale: Evaluating Reasoning Traces as Retrieval Sources for Domain Adaptation of Small Language Models
- BIBERT-Pipe on Biomedical Nested Named Entity Linking at BioASQ 2025
- ReProCon: Scalable and Resource-Efficient Few-Shot Biomedical Named Entity Recognition
- A Common Pipeline for Harmonizing Electronic Health Record Data for Translational Research
- A Survey of Long-Document Retrieval in the PLM and LLM Era
- BALI: Enhancing Biomedical Language Representations through Knowledge Graph and Language Model Alignment
- Automated Hierarchical Graph Construction for Multi-source Electronic Health Records
- Semantic Analysis of SNOMED CT Concept Co-occurrences in Clinical Documentation using MIMIC-IV
- DSRAG: A Domain-Specific Retrieval Framework Based on Document-derived Multimodal Knowledge Graph
- OPRA-Vis: Visual Analytics System to Assist Organization-Public Relationship Assessment with Large Language Models
- Weakly Supervised Medical Entity Extraction and Linking for Chief Complaints
- Unified Supervision For Vision-Language Modeling in 3D Computed Tomography
- Cross-lingual Natural Language Processing on Limited Annotated Case/Radiology Reports in English and Japanese: Insights from the Real-MedNLP Workshop
- Benchmarking GPT-5 in Radiation Oncology: Measurable Gains, but Persistent Need for Expert Oversight
- Overview of BioASQ 2025: The Thirteenth BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Question Answering
- Overview of BioASQ 2024: The twelfth BioASQ challenge on Large-Scale Biomedical Semantic Indexing and Question Answering
- Benchmarking GPT-5 for biomedical natural language processing
- Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
- BERT might be Overkill: A Tiny but Effective Biomedical Entity Linker based on Residual Convolutional Neural Networks
- Scalable Scientific Interest Profiling Using Large Language Models
- A Language-Signal-Vision Multimodal Framework for Multitask Cardiac Analysis
- Extracting Post-Acute Sequelae of SARS-CoV-2 Infection Symptoms from Clinical Notes via Hybrid Natural Language Processing
- Dataset Construction for Training LLM to Learn Analog Circuit Knowledge
- Specialised or Generic? Tokenization Choices for Radiology Language Models
- Benchmarking the Medical Understanding and Reasoning of Large Language Models in Arabic Healthcare Tasks
- Improving BERT Model Using Contrastive Learning for Biomedical Relation Extraction
- A semantic atlas of journals: Structure, position, and dispersion
- MedSumGraph: enhancing GraphRAG for medical QA with summarization and optimized prompts
- Graph-Augmented Language Model Framework for Health Misinformation Detection
- Explanatory argument extraction of correct answers in resident medical exams
- On the effectiveness of multimodal privileged knowledge distillation in two vision transformer based diagnostic applications
- Large Language Model's Multi-Capability Alignment in Biomedical Domain
- RooseBERT: A New Deal For Political Language Modelling
- CoughViT: A Self-Supervised Vision Transformer for Cough Audio Representation Learning
- Welcome New Doctor: Continual Learning with Expert Consultation and Autoregressive Inference for Whole Slide Image Analysis
- ChEmbed: Enhancing Chemical Literature Search Through Domain-Specific Text Embeddings
- OpenMed NER: Open-Source, Domain-Adapted State-of-the-Art Transformers for Biomedical NER Across 12 Public Datasets
- Harnessing Collective Intelligence of LLMs for Robust Biomedical QA: A Multi-Model Approach
- Towards Efficient Medical Reasoning with Minimal Fine-Tuning Data
- CancerBERT: a BERT model for Extracting Breast Cancer Phenotypes from Electronic Health Records
- How Far Are AI Scientists from Changing the World?
- Read, Attend, and Code: Pushing the Limits of Medical Codes Prediction from Clinical Notes by Machines
- Cardiac-CLIP: A Vision-Language Foundation Model for 3D Cardiac CT Images
Related