Making Pre-trained Language Models Better Few-shot Learners
2020/12/31 by Tianyu Gao, Adam Fisch, Gao, Tianyu +3 · 116 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2012.15723
Accepted to ACL 2021. The code is publicly available at https://github.com/princeton-nlp/LM-BFF
openalex publication_date 2020/12/31 · arxiv created 2021/06/02 · arxiv updated 2021/06/03 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
The recent GPT-3 model (Brown et al., 2020) achieves remarkable few-shot performance solely by leveraging a natural-language prompt and a few task demonstrations as input context. Inspired by their findings, we study few-shot learning in a more practical scenario, where we use smaller language models for which fine-tuning is computationally efficient. We present LM-BFF--better few-shot fine-tuning of language models--a suite of simple and complementary techniques for fine-tuning language models on a small number of annotated examples. Our approach includes (1) prompt-based fine-tuning together with a novel pipeline for automating prompt generation; and (2) a refined strategy for dynamically and selectively incorporating demonstrations into each context. Finally, we present a systematic evaluation for analyzing few-shot performance on a range of NLP tasks, including classification and regression. Our experiments demonstrate that our methods combine to dramatically outperform standard fine-tuning procedures in this low resource setting, achieving up to 30% absolute improvement, and 11% on average across all tasks. Our approach makes minimal assumptions on task resources and domain expertise, and hence constitutes a strong task-agnostic method for few-shot learning.
Citations
Cited by
- Merge before Forget: A Single LoRA Continual Learning via Continual Merging
- Imprompt: A Language Framework for Prompt Programming
- Retrieval-augmented Prompt Learning for Pre-trained Foundation Models
- An Exploratory Study of Bayesian Prompt Optimization for Test-Driven Code Generation with Large Language Models
- How Prompts Move Language Model Behavior: Frames, Salience, and Construal as Semantic Control
- Rethinking Label Consistency of In-Context Learning: An Implicit Transductive Label Propagation Perspective
- LabelFusion: Learning to Fuse LLMs and Transformer Classifiers for Robust Text Classification
- Do Generalisation Results Generalise?
- Efficient Text Classification with Conformal In-Context Learning
- Empirical Prompt Engineering for Construct Identification with Large Language Models
- 5G Network Automation Using Local Large Language Models and Retrieval-Augmented Generation
- Semantic Anchors in In-Context Learning: Why Small LLMs Cannot Flip Their Labels
- From Words to Wisdom: Discourse Annotation and Baseline Models for Student Dialogue Understanding
- Think First, Assign Next (ThiFAN-VQA): A Two-stage Chain-of-Thought Framework for Post-Disaster Damage Assessment
- Zero-Shot Grammar Competency Estimation Using Large Language Model Generated Pseudo Labels
- On-Device Fine-Tuning via Backprop-Free Zeroth-Order Optimization
- Structured Definitions and Segmentations for Legal Reasoning in LLMs: A Study on Indian Legal Data
- GFT: Graph Feature Tuning for Efficient Point Cloud Analysis
- Multi-refined Feature Enhanced Sentiment Analysis Using Contextual Instruction
- Learning to Prompt for Vision-Language Models
- Stay Tuned: Improving Sentiment Analysis and Stance Detection Using Large Language Models
- P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks
- Semantic-Preserving Cross-Style Visual Reasoning for Robust Multi-Modal Understanding in Large Vision-Language Models
- Large Language Models Meet Text-Attributed Graphs: A Survey of Integration Frameworks and Applications
- SODBench: A Large Language Model Approach to Documenting Spreadsheet Operations
- All You Need is One: Capsule Prompt Tuning with a Single Vector
- Controlling the Risk of Corrupted Contexts for Language Models via Early-Exiting
- MedREK: Retrieval-Based Editing for Medical LLMs with Key-Aware Prompts
- A Survey on Evaluation of Large Language Models
- Prompts Generalize with Low Data: Non-vacuous Generalization Bounds for Optimizing Prompts with More Informative Priors
- Evolutionary Computation as Natural Generative AI
- Security and Privacy Challenges of Large Language Models: A Survey
- Differentiable Prompt Makes Pre-trained Language Models Better Few-shot Learners
- A Retail-Corpus for Aspect-Based Sentiment Analysis with Large Language Models
- Evolution of Concepts in Language Model Pre-Training
- Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs
- The Few-shot Dilemma: Over-prompting Large Language Models
- NeuroStrike: Neuron-Level Attacks on Aligned LLMs
- Graph-Enhanced Retrieval-Augmented Question Answering for E-Commerce Customer Support
- Topic Coverage-based Demonstration Retrieval for In-Context Learning
- Are Humans as Brittle as Large Language Models?
- Fuzzy, Symbolic, and Contextual: Enhancing LLM Instruction via Cognitive Scaffolding
- Membership Inference Attacks on LLM-based Recommender Systems
- TrackRec: Iterative Alternating Feedback with Chain-of-Thought via Preference Alignment for Recommendation
- Label Verbalization and Entailment for Effective Zero- and Few-Shot Relation Extraction
- When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs
- STEP: Stepwise Curriculum Learning for Context-Knowledge Fusion in Conversational Recommendation
- Prompt-Based Approach for Czech Sentiment Analysis
- CATP: Contextually Adaptive Token Pruning for Efficient and Enhanced Multimodal In-Context Learning
- Prompt Tuning for Few-Shot Continual Learning Named Entity Recognition
- A Fuzzy Logic Prompting Framework for Large Language Models in Adaptive and Uncertain Tasks
- ETTA: Efficient Test-Time Adaptation for Vision-Language Models through Dynamic Embedding Updates
- GeoSR: Cognitive-Agentic Framework for Probing Geospatial Knowledge Boundaries via Iterative Self-Refinement
- Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D Generation
- Reservoir Computing as a Language Model
- Few-Sample Named Entity Recognition for Security Vulnerability Reports by Fine-Tuning Pre-Trained Language Models
- Swin-TUNA : A Novel PEFT Approach for Accurate Food Image Segmentation
- FedDPG: An Adaptive Yet Efficient Prompt-tuning Approach in Federated Learning Settings
- MERA Code: A Unified Framework for Evaluating Code Generation Across Tasks
- BERT is to NLP what AlexNet is to CV: Can Pre-Trained Language Models Identify Analogies?
- Calibrate Before Use: Improving Few-Shot Performance of Language Models
- LFPT5: A Unified Framework for Lifelong Few-shot Language Learning Based on Prompt Tuning of T5
- What Should LLMs Forget? Quantifying Personal Data in LLMs for Right-to-Be-Forgotten Requests
- Factual Probing Is [MASK]: Learning vs. Learning to Recall
- Tiny Reward Models
- Stabilizing Black-Box Prompt Optimization with Textual Regularization and Signal Aggregation
- Knowledgeable or Educated Guess? Revisiting Language Models as Knowledge Bases
- Your Pretrained Model Tells the Difficulty Itself: A Self-Adaptive Curriculum Learning Paradigm for Natural Language Understanding
- Adversarial Demonstration Learning for Low-resource NER Using Dual Similarity
- Improving Robustness of Foundation Models in Domain Adaptation with Soup-Adapters
- Unveiling Effective In-Context Configurations for Image Captioning: An External & Internal Analysis
- Dual Modality-Aware Gated Prompt Tuning for Few-Shot Multimodal Sarcasm Detection
- Do Language Models Perform Generalizable Commonsense Inference?
- Few-Shot Bot: Prompt-Based Learning for Dialogue Systems
- Modeling Data Diversity for Joint Instance and Verbalizer Selection in Cold-Start Scenarios
- VisualPrompter: Prompt Optimization with Visual Feedback for Text-to-Image Synthesis
- Theoretical Modeling of Large Language Model Self-Improvement Training Dynamics Through Solver-Verifier Gap
- Memory Savings at What Cost? A Study of Alternatives to Backpropagation
- Structuralist Approach to AI Literary Criticism: Leveraging Greimas Semiotic Square for Large Language Models
- Multi-Objective Recommendation in the Era of Generative AI: A Survey of Recent Progress and Future Prospects
- Efficient and Privacy-Preserving Soft Prompt Transfer for LLMs
- RiOT: Efficient Prompt Refinement with Residual Optimization Tree
- Causally Steered Diffusion for Automated Video Counterfactual Generation
- Neural Network Reprogrammability: A Unified Theme on Model Reprogramming, Prompt Tuning, and Prompt Instruction
- Prompt Candidates, then Distill: A Teacher-Student Framework for LLM-driven Data Annotation
- Hierarchical Self-Prompting SAM: A Prompt-Free Medical Image Segmentation Framework
- Statement-Tuning Enables Efficient Cross-lingual Generalization in Encoder-only Models
- Few-Shot Optimization for Sensor Data Using Large Language Models: A Case Study on Fatigue Detection
- Optimizing LLM-Based Multi-Agent System with Textual Feedback: A Case Study on Software Development
- Only Large Weights (And Not Skip Connections) Can Prevent the Perils of Rank Collapse
- TACO: Enhancing Multimodal In-context Learning via Task Mapping-Guided Sequence Configuration
- CAMA: Enhancing Multimodal In-Context Learning with Context-Aware Modulated Attention
- Distribution Prompting: Understanding the Expressivity of Language Models Through the Next-Token Distributions They Can Produce
- Bridging Generative and Discriminative Learning: Few-Shot Relation Extraction via Two-Stage Knowledge-Guided Pre-training
- Fast RoPE Attention: Combining the Polynomial Method and Fast Fourier Transform
- DesignFromX: Empowering Consumer-Driven Design Space Exploration through Feature Composition of Referenced Products
- The Hitchhikers Guide to Production-ready Trustworthy Foundation Model powered Software (FMware)
- PLHF: Prompt Optimization with Few-Shot Human Feedback
- ConQX: Semantic Expansion of Spoken Queries for Intent Detection based on Conditioned Text Generation
- Towards Embodiment Scaling Laws in Robot Locomotion
- Template-Based Named Entity Recognition Using BART
- Retrieval-Enhanced Few-Shot Prompting for Speech Event Extraction
- Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact
- Perturbation-efficient Zeroth-order Optimization for Hardware-friendly On-device Training
- Coreference Resolution for Vietnamese Narrative Texts
- ContiguousKV: Accelerating LLM Prefill with Granularity-Aligned KV Cache Management
- AdaptRec: A Self-Adaptive Framework for Sequential Recommendations with Large Language Models
- Few-shot Hate Speech Detection Based on the MindSpore Framework
- CLIP-Powered Domain Generalization and Domain Adaptation: A Comprehensive Survey
- A Comprehensive Survey of Challenges and Opportunities of Few-Shot Learning Across Multiple Domains
- Knowledge Acquisition on Mass-shooting Events via LLMs for AI-Driven Justice
- Adaptive Additive Parameter Updates of Vision Transformers for Few-Shot Continual Learning
- Mimic In-Context Learning for Multimodal Tasks
- S2R-HDR: A Large-Scale Rendered Dataset for HDR Fusion
- AI as a resource for the clarification of medical terminology
- System Log Parsing with Large Language Models: A Review
Related