Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?
2022/02/25 by Sewon Min, Min, Sewon, Xinxi Lyu +10 · 116 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2202.12837
openalex publication_date 2022/02/25 · openalex created_date 2022/04/03 · openalex updated_date 2026/07/28
Abstract
Large language models (LMs) are able to in-context learn -- perform a new task via inference alone by conditioning on a few input-label pairs (demonstrations) and making predictions for new inputs. However, there has been little understanding of how the model learns and which aspects of the demonstrations contribute to end task performance. In this paper, we show that ground truth demonstrations are in fact not required -- randomly replacing labels in the demonstrations barely hurts performance on a range of classification and multi-choce tasks, consistently over 12 different models including GPT-3. Instead, we find that other aspects of the demonstrations are the key drivers of end task performance, including the fact that they provide a few examples of (1) the label space, (2) the distribution of the input text, and (3) the overall format of the sequence. Together, our analysis provides a new way of understanding how and why in-context learning works, while opening up new questions about how much can be learned from large language models through inference alone.
Cited by
- Single LLM Debate, MoLaCE: Mixture of Latent Concept Experts Against Confirmation Bias
- Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex
- Where Steering Signals Come From: Activation Source Selection in Activation Steering
- Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe
- When Does Few-Shot Prompting Help? A Systematic Empirical Study of Shot-Count Effects Across Model Scale, Architecture, and Output Parsing Robustness
- Beyond Exact Match: How Evaluation Methodology Dominates Model Choice in LLM-Based Product Attribute Extraction
- LoFT-LLM: Low-Frequency Time-Series Forecasting with Large Language Models
- From Shortcut to Induction Head: How Data Diversity Shapes Algorithm Selection in Transformers
- DACE For Railway Acronym Disambiguation
- Task Schema and Binding: A Double Dissociation Study of In-Context Learning
- Case Prompting to Mitigate Large Language Model Bias for ICU Mortality Prediction
- Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based Perspective
- FiNERweb: Datasets and Artifacts for Scalable Multilingual Named Entity Recognition
- Textual Gradients are a Flawed Metaphor for Automatic Prompt Optimization
- Understanding Structured Financial Data with LLMs: A Case Study on Fraud Detection
- How Prompts Move Language Model Behavior: Frames, Salience, and Construal as Semantic Control
- Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors
- To Think or Not to Think: The Hidden Cost of Meta-Training with Excessive CoT Examples
- Attention as Binding: A Vector-Symbolic Perspective on Transformer Reasoning
- Exploiting the Randomness of Large Language Models (LLM) in Text Classification Tasks: Locating Privileged Documents in Legal Matters
- The Road of Adaptive AI for Precision in Cybersecurity
- Semantic Anchors in In-Context Learning: Why Small LLMs Cannot Flip Their Labels
- Learning Multi-Access Point Coordination in Agentic AI Wi-Fi with Large Language Models
- Semantics as a Shield: Label Disguise Defense (LDD) against Prompt Injection in LLM Sentiment Classification
- SDA: Steering-Driven Distribution Alignment for Open LLMs without Fine-Tuning
- Mitigating Label Length Bias in Large Language Models
- RAG-Driven Data Quality Governance for Enterprise ERP Systems
- Grounded by Experience: Generative Healthcare Prediction Augmented with Hierarchical Agentic Retrieval
- Agent READMEs: An Empirical Study of Context Files for Agentic Coding
- Generative Caching for Structurally Similar Prompts and Responses
- Adaptive Multi-Agent Response Refinement in Conversational Systems
- CG-TTRL: Context-Guided Test-Time Reinforcement Learning for On-Device Large Language Models
- DRAGON: Guard LLM Unlearning in Context via Negative Detection and Reasoning
- ReleaseEval: A Benchmark for Evaluating Language Models in Automated Release Note Generation
- MARS-SQL: A multi-agent reinforcement learning framework for Text-to-SQL
- CodeAlignBench: Assessing Code Generation Models on Developer-Preferred Code Adjustments
- Predicate Renaming via Large Language Models
- Prompt Framing Distorts Count Based Evaluation of LLM Error Detection: Evidence from Numeric Anchoring
- EsoLang-Bench: Evaluating Genuine Reasoning in Large Language Models via Esoteric Programming Languages
- From Reviews to Actionable Insights: An LLM-Based Approach for Attribute and Feature Extraction
- How Context Shapes Truth: Geometric Transformations of Statement-level Truth Representations in LLMs
- LAMUS: A Large-Scale Corpus for Legal Argument Mining from U.S. Caselaw using LLMs
- Beyond One-Size-Fits-All: Personalized Harmful Content Detection with In-Context Learning
- MoS-VLA: A Vision-Language-Action Model with One-Shot Skill Adaptation
- Can Language Models Compose Skills In-Context?
- R2ComSync: Improving Code-Comment Synchronization with In-Context Learning and Reranking
- Preliminary Use of Vision Language Model Driven Extraction of Mouse Behavior Towards Understanding Fear Expression
- Beyond More Context: Retrieval Diversity Boosts Multi-Turn Intent Understanding
- LLM-ERM: Sample-Efficient Program Learning via LLM-Guided Search
- Schema for In-Context Learning
- Reasoning Pattern Matters: Learning to Reason without Human Rationales
- In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning
- Training-Free In-Context Forensic Chain for Image Manipulation Detection and Localization
- Logit Arithmetic Elicits Long Reasoning Capabilities Without Training
- Repairing Regex Vulnerabilities via Localization-Guided Instructions
- Graph Diffusion Transformers are In-Context Molecular Designers
- On the Relationship Between the Choice of Representation and In-Context Learning
- MHA-RAG: Improving Efficiency, Accuracy, and Consistency by Encoding Exemplars as Soft Prompts
- GRAD: Generative Retrieval-Aligned Demonstration Sampler for Efficient Few-Shot Reasoning
- BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
- Talk in Pieces, See in Whole: Disentangling and Hierarchical Aggregating Representations for Language-based Object Detection
- Localizing Task Recognition and Task Learning in In-Context Learning via Attention Head Analysis
- ReasonCACHE: Teaching LLMs To Reason Without Weight Updates
- The Impact of Role Design in In-Context Learning for Large Language Models
- StateX: Enhancing RNN Recall via Post-training State Expansion
- Context Parametrization with Compositional Adapters
- Induction Signatures Are Not Enough: A Matched-Compute Study of Load-Bearing Structure in In-Context Learning
- A Structured Knowledge Infrastructure for Domain-Specific Data Asset Discovery
- Are Smaller Open-Weight LLMs Closing the Gap to Proprietary Models for Biomedical Question Answering?
- Uncovering Implicit Bias in Large Language Models with Concept Learning Dataset
- KITE: Kernelized and Information Theoretic Exemplars for In-Context Learning
- SciEvent: Benchmarking Multi-domain Scientific Event Extraction
- On the Use of Agentic Coding Manifests: An Empirical Study of Claude Code
- TextMineX: Data, Evaluation Framework and Ontology-guided LLM Pipeline for Humanitarian Mine Action
- Latent Traits and Cross-Task Transfer: Deconstructing Dataset Interactions in LLM Fine-tuning
- RAGs to Riches: RAG-like Few-shot Learning for Large Language Model Role-playing
- Public Data Assisted Differentially Private In-Context Learning
- Reasoning Under Uncertainty: Exploring Probabilistic Reasoning Capabilities of LLMs
- Test-Time Warmup for Multimodal Large Language Models
- InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning
- Abduct, Act, Predict: Scaffolding Causal Inference for Automated Failure Attribution in Multi-Agent Systems
- Fluent but Unfeeling: The Emotional Blind Spots of Language Models
- ALLabel: Three-stage Active Learning for LLM-based Entity Recognition using Demonstration Retrieval
- In-Context Learning Enhanced Credibility Transformer
- MachineLearningLM: Scaling Many-shot In-context Learning via Continued Pretraining
- Do LLMs Adhere to Label Definitions? Examining Their Receptivity to External Label Definitions
- Retrieval Enhanced Feedback via In-context Neural Error-book
- Generative KI für TA
- Extracting OPQRST in Electronic Health Records using Large Language Models with Reasoning
- Baichuan-M2: Scaling Medical Capability with Large Verifier System
- Rethinking the Chain-of-Thought: The Roles of In-Context Learning and Pre-trained Priors
- Supervised In-Context Fine-Tuning for Generative Sequence Labeling
- Just-in-time and distributed task representations in language models
- InSQuAD: In-Context Learning for Efficient Retrieval via Submodular Mutual Information to Enforce Quality and Diversity
- CAPE: Context-Aware Personality Evaluation Framework for Large Language Models
- Linear-Time Demonstration Selection for In-Context Learning via Gradient Estimation
- Named Entity Recognition of Historical Text via Large Language Model
- Self-Guided Function Calling in Large Language Models via Stepwise Experience Recall
- Adversarial Attacks against Neural Ranking Models via In-Context Learning
- In-Context Iterative Policy Improvement for Dynamic Manipulation
- Are Large Pre-trained Vision Language Models Effective Construction Safety Inspectors?
- REFN: A Reinforcement-Learning-From-Network Framework against 1-day/n-day Exploitations
- A Survey on Training-free Alignment of Large Language Models
- IROTE: Human-like Traits Elicitation of Large Language Model via In-Context Self-Reflective Optimization
- ColorGPT: Leveraging Large Language Models for Multimodal Color Recommendation
- In-situ Value-aligned Human-Robot Interactions with Physical Constraints
- AttriLens-Mol: Attribute Guided Reinforcement Learning for Molecular Property Prediction with Large Language Models
- Beyond Meme Templates: Limitations of Visual Similarity Measures in Meme Matching
- LECTOR: LLM-Enhanced Concept-based Test-Oriented Repetition for Adaptive Spaced Learning
- Can LLMs Generate High-Quality Task-Specific Conversations?
- To Theoretically Understand Transformer-Based In-Context Learning for Optimizing CSMA
- DICE: Dynamic In-Context Example Selection in LLM Agents via Efficient Knowledge Transfer
- Where to show Demos in Your Prompt: A Positional Bias of In-Context Learning
- MapAgent: Trajectory-Constructed Memory-Augmented Planning for Mobile Task Automation
- Can large language models assist choice modelling? Insights into prompting strategies and current models capabilities
- StyleAdaptedLM: Enhancing Instruction Following Models with Efficient Stylistic Transfer
Related