Don't Stop Pretraining: Adapt Language Models to Domains and Tasks
2020/04/23 by Gururangan, Suchin, Marasović, Ana, Swayamdipta, Swabha +4 · 103 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
paper · doi:10.48550/arxiv.2004.10964
Abstract
Language models pretrained on text from a wide variety of sources form the foundation of today's NLP. In light of the success of these broad-coverage models, we investigate whether it is still helpful to tailor a pretrained model to the domain of a target task. We present a study across four domains (biomedical and computer science publications, news, and reviews) and eight classification tasks, showing that a second phase of pretraining in-domain (domain-adaptive pretraining) leads to performance gains, under both high- and low-resource settings. Moreover, adapting to the task's unlabeled data (task-adaptive pretraining) improves performance even after domain-adaptive pretraining. Finally, we show that adapting to a task corpus augmented using simple data selection strategies is an effective alternative, especially when resources for domain-adaptive pretraining might be unavailable. Overall, we consistently find that multi-phase adaptive pretraining offers large gains in task performance.
Cited by
- From transcription to semantic corpus analysis: unsupervised learning of sentence representations for ancient languages
- Fake News Classification in Urdu: A Domain Adaptation Approach for a Low-Resource Language
- Agentic Physical AI toward a Domain-Specific Foundation Model for Energy Systems: A Case Study on Nuclear Reactor Control
- seqLens: Optimizing Language Models for Genomic Predictions
- Adapter Merging Reactivates Latent Reasoning Traces: A Mechanism Analysis
- Alchemist: Unlocking Efficiency in Text-to-Image Model Training via Meta-Gradient Data Selection
- ReactorFold: Generative discovery of nuclear reactor cores via emergent physical reasoning
- Mining Legal Arguments to Study Judicial Formalism
- Zero-shot 3D Map Generation with LLM Agents: A Dual-Agent Architecture for Procedural Content Generation
- System Report for CCL25-Eval Task 10: Prompt-Driven Large Language Model Merge for Fine-Grained Chinese Hate Speech Detection
- Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages
- The Road of Adaptive AI for Precision in Cybersecurity
- Gradient Descent with Provably Tuned Learning-rate Schedules
- Adapting Large Language Models to Low-Resource Tibetan: A Two-Stage Continual and Supervised Fine-Tuning Study
- Network Self-Configuration based on Fine-Tuned Small Language Models
- Comparative Analysis of 47 Context-Based Question Answer Models Across 8 Diverse Datasets
- MortgageLLM: Domain-Adaptive Pretraining with Residual Instruction Transfer, Alignment Tuning, and Task-Specific Routing
- Building Domain-Specific Small Language Models via Guided Data Generation
- Bridging VLMs and Embodied Intelligence with Deliberate Practice Policy Optimization
- Tokenize Once, Recommend Anywhere: Unified Item Tokenization for Multi-domain LLM-based Recommendation
- NeuroLex: A Lightweight Domain Language Model for EEG Report Understanding and Generation
- Evaluating the Ability of Large Language Models to Identify Adherence to CONSORT Reporting Guidelines in Randomized Controlled Trials: A Methodological Evaluation Study
- BIRD: Bronze Inscription Restoration and Dating
- Classification of Hope in Textual Data using Transformer-Based Models
- Concept-Based Interpretability for Toxicity Detection
- PriVi: Towards A General-Purpose Video Model For Primate Behavior In The Wild
- Kunlun Anomaly Troubleshooter: Enabling Kernel-Level Anomaly Detection and Causal Reasoning for Large Model Distributed Inference
- ManufactuBERT: Efficient Continual Pretraining for Manufacturing
- MIDI-LLM: Adapting Large Language Models for Text-to-MIDI Music Generation
- Exploring and Mitigating Gender Bias in Encoder-Based Transformer Models
- Multilingual BERT language model for medical tasks: Evaluation on domain-specific adaptation and cross-linguality
- From Amateur to Master: Infusing Knowledge into LLMs via Automated Curriculum Learning
- Layer of Truth: Probing Belief Shifts under Continual Pre-Training Poisoning
- Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges
- Extractive versus Generative Language Models for Political Conflict Text Classification
- Beyond One-Size-Fits-All: Personalized Harmful Content Detection with In-Context Learning
- Network Intrusion Detection: Evolution from Conventional Approaches to LLM Collaboration and Emerging Risks
- A Survey on LLM Mid-Training
- Robust Uncertainty Quantification for Self-Evolving Large Language Models via Continual Domain Pretraining
- Generating Auxiliary Tasks with Reinforcement Learning
- Low-Resource Dialect Adaptation of Large Language Models: A French Dialect Case-Study
- From Slides to Chatbots: Enhancing Large Language Models with University Course Materials
- PatenTEB: A Comprehensive Benchmark and Model Family for Patent Text Embedding
- VESSA: Video-based objEct-centric Self-Supervised Adaptation for Visual Foundation Models
- IKnow: Instruction-Knowledge-Aware Continual Pretraining for Effective Domain Adaptation
- Adapting Multilingual Models to Code-Mixed Tasks via Model Merging
- AtlasKV: Augmenting LLMs with Billion-Scale Knowledge Graphs in 20GB VRAM
- Qomhra: A Bilingual Irish and English Large Language Model
- Midtraining Bridges Pretraining and Posttraining Distributions
- Cognitive-Aligned Spatio-Temporal Large Language Models For Next Point-of-Interest Prediction
- First Attentions Last: Better Exploiting First Attentions for Efficient Transformer Training
- Rewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model Behavior
- Sparse Subnetwork Enhancement for Underrepresented Languages in Large Language Models
- A-IPO: Adaptive Intent-driven Preference Optimization
- Autonomous Agents for Scientific Discovery: Orchestrating Scientists, Language, Code, and Physics
- Exploring Cross-Lingual Knowledge Transfer via Transliteration-Based MLM Fine-Tuning for Critically Low-resource Chakma Language
- Understanding the Effects of Domain Finetuning on LLMs
- DACIP-RC: Domain Adaptive Continual Instruction Pre-Training via Reading Comprehension on Business Conversations
- SliceFine: The Universal Winning-Slice Hypothesis for Pretrained Networks
- Beyond Monolingual Assumptions: A Survey of Code-Switched NLP in the Era of Large Language Models
- DACP: Domain-Adaptive Continual Pre-Training of Large Language Models for Phone Conversation Summarization
- Contrastive Learning Using Graph Embeddings for Domain Adaptation of Language Models in the Process Industry
- AWARE, Beyond Sentence Boundaries: A Contextual Transformer Framework for Identifying Cultural Capital in STEM Narratives
- Train on Validation (ToV): Fast data selection with applications to fine-tuning
- Quantifying Semantic Shift in Financial NLP: Robust Metrics for Market Prediction Stability
- CustomIR: Unsupervised Fine-Tuning of Dense Embeddings for Known Document Corpora
- Predicting Training Re-evaluation Curves Enables Effective Data Curriculums for LLMs
- Metaphor identification using large language models: A comparison of RAG, prompt engineering, and fine-tuning
- WirelessMathLM: Teaching Mathematical Reasoning for LLMs in Wireless Communications with Reinforcement Learning
- Learning the Wrong Lessons: Syntactic-Domain Spurious Correlations in Language Models
- Policy Compatible Skill Incremental Learning via Lazy Learning Interface
- Embedding Domain Knowledge for Large Language Models via Reinforcement Learning from Augmented Generation
- Detoxifying Large Language Models via Autoregressive Reward Guided Representation Editing
- Improving Mental Health Screening and Early Risk Detection in Spanish
- Beyond Rephrasing: Book-Level Organization Improves Synthetic Textbook Data for Mid-Training
- Memory in Large Language Models: Mechanisms, Evaluation and Evolution
- Visual Instruction Pretraining for Domain-Specific Foundation Models
- PG-CE: A Progressive Generation Dataset with Constraint Enhancement for Controllable Text Generation
- Domain-Adaptive Pre-Training for Arabic Aspect-Based Sentiment Analysis: A Comparative Study of Domain Adaptation and Fine-Tuning Strategies
- Rethinking the Role of Text Complexity in Language Model Pretraining
- Optimizing Product Deduplication in E-Commerce with Multimodal Embeddings
- Deep learning and abstractive summarisation for radiological reports: an empirical study for adapting the PEGASUS models' family with scarce data
- Active Domain Knowledge Acquisition with 100-Dollar Budget: Enhancing LLMs via Cost-Efficient, Expert-Involved Interaction in Sensitive Domains
- Boosting Data Utilization for Multilingual Dense Retrieval
- Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector
- Augmented Fine-Tuned LLMs for Enhanced Recruitment Automation
- ChatGPT-generated texts show authorship traits that identify them as non-human
- Hierarchical Section Matching Prediction (HSMP) BERT for Fine-Grained Extraction of Structured Data from Hebrew Free-Text Radiology Reports in Crohn's Disease
- Linear-Time Demonstration Selection for In-Context Learning via Gradient Estimation
- Calibrated Semantic Diffusion: A p-Laplacian Synthesis with Learnable Dissipation, Quantified Constants, and Graph-Aware Calibration
- LegalΔ: Enhancing Legal Reasoning in LLMs via Reinforcement Learning with Chain-of-Thought Guided Information Gain
- When Does Language Transfer Help? Sequential Fine-Tuning for Cross-Lingual Euphemism Detection
- Dataset Construction for Training LLM to Learn Analog Circuit Knowledge
- ALAS: Autonomous Learning Agent for Self-Updating Language Models
- Memory Decoder: A Pretrained, Plug-and-Play Memory for Large Language Models
- Effortless Vision-Language Model Specialization in Histopathology without Annotation
- Arce: Augmented Roberta with Contextualized Elucidations for Ner in Automated Rule Checking
- Comparing Knowledge Injection Methods for LLMs in a Low-Resource Regime
- Sensitivity of Stability: Theoretical & Empirical Analysis of Replicability for Adaptive Data Selection in Transfer Learning
- Multidimensional classification of posts for online course discussion forum curation
- LLM-based IR-system for Bank Supervisors
- OpenMed NER: Open-Source, Domain-Adapted State-of-the-Art Transformers for Biomedical NER Across 12 Public Datasets
- Measuring Time-Series Dataset Similarity using Wasserstein Distance
Related