Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
2023/10/17 by Melanie Sclar, Sclar, Melanie, Yejin Choi +5 · 4 voices · 138 citations
Computer Science · Engineering · #Ferroelectric and Negative Capacitance Devices #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2310.11324
openalex publication_date 2023/10/17 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
As large language models (LLMs) are adopted as a fundamental component of language technologies, it is crucial to accurately characterize their performance. Because choices in prompt design can strongly influence model behavior, this design process is critical in effectively using any modern pre-trained generative language model. In this work, we focus on LLM sensitivity to a quintessential class of meaning-preserving design choices: prompt formatting. We find that several widely used open-source LLMs are extremely sensitive to subtle changes in prompt formatting in few-shot settings, with performance differences of up to 76 accuracy points when evaluated using LLaMA-2-13B. Sensitivity remains even when increasing model size, the number of few-shot examples, or performing instruction tuning. Our analysis suggests that work evaluating LLMs with prompting-based methods would benefit from reporting a range of performance across plausible prompt formats, instead of the currently-standard practice of reporting performance on a single format. We also show that format performance only weakly correlates between models, which puts into question the methodological validity of comparing models with an arbitrarily chosen, fixed prompt format. To facilitate systematic analysis we propose FormatSpread, an algorithm that rapidly evaluates a sampled set of plausible prompt formats for a given task, and reports the interval of expected performance without accessing model weights. Furthermore, we present a suite of analyses that characterize the nature of this sensitivity, including exploring the influence of particular atomic perturbations and the internal representation of particular formats.
Cited by
- Grounding latent algorithm routing in transformer reasoning
- The Severance Problem: LLMs are Unaware of the Person Beyond the Prompt
- Instrument Effects in Language-Model Honesty Evaluation: An Auditable Single-System Demonstration
- LaSEr-Edit: Localized Span-level Error Editing with Energy-based Localization
- Reference-Free Evaluation of Reasoning in Open-Ended Question Answering
- A Systematic Evaluation of Traditional Privacy Policy Analysis Tools Against LLMs
- Prompt Robustness Is Task-Dependent: Comparing Objective and Belief-Style Questions in LLM Evaluation
- Prompt Design at Scale: How Format, Instruction Count, and Context Length Shape Instruction Adherence and Hallucination in Large Language Models
- Structured Output Collapses Answer Diversity Across 44 Language Models
- Bridging the Information Gap: Semantic Densification and Hindsight Distillation for Cold-Start Prediction
- When Perplexity Lies: Generation-Focused Distillation of Hybrid Sequence Models
- Validating LLMs in social science: Epistemic threats and emerging norms
- The Riddle Riddle: Testing Flexible Reasoning in Large Language Models and Humans
- Self-Harness: Harnesses That Improve Themselves
- A Knowledge-Injection Framework for Zero-Shot Adaptation of LLMs to Delirium Prediction
- Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation
- How Open Must Language Models be to Enable Reliable Scientific Inference?
- This Treatment Works, Right? Evaluating LLM Sensitivity to Patient Question Framing in Medical QA
- Mitigating Conversational Inertia in Multi-Turn Agents
- RegCheck: A tool for structured comparisons between study registrations and papers
- Benchmarks Saturate When The Model Gets Smarter Than The Judge
- Mind Your Tone: Investigating How Prompt Politeness Affects LLM Accuracy (short paper)
- What is a protest anyway? Codebook conceptualization is still a first-order concern in LLM-era classification
- A Single Character can Make or Break Your LLM Evals
- The threat of analytic flexibility in using large language models to simulate human data
- Is In-Context Learning Learning?
- Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
- Many AI Analysts, One Dataset: Navigating the Agentic Data Science Multiverse
- The Impact of Prompt Programming on Function-Level Code Generation
- Are Prompt Optimizers Blind? Cross-Modal Visual Feedback for Automatic Prompt Optimization
- Measuring and Improving Behavioral Consistency in Large Language Models through Fact-Heuristic-Emotion State Enforcement
- Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy
- CARE-MH: Towards Unified, Reproducible, and Comparable Evaluation of Mental Health LLMs
- Statistical realism is not evidence that LLMs can estimate treatment effects in social science experiments
- The Cartesian Cut in Agentic AI
- Algorithmic Blindness in Large Language Models: A Calibration Study of Performance Prediction
- LLMs on Drugs: Language Models Are Few-Shot Consumers
- Visually Prompted Benchmarks Are Surprisingly Fragile
- When the Gold Standard Isn't Necessarily Standard: Challenges of Evaluating the Translation of User-Generated Content
- Revisiting the Reliability of Language Models in Instruction-Following
- How Prompts Move Language Model Behavior: Frames, Salience, and Construal as Semantic Control
- Chasing Shadows: Pitfalls in LLM Security Research
- QSTN: A Modular Framework for Robust Questionnaire Inference with Large Language Models
- Mary, the Cheeseburger-Eating Vegetarian: Do LLMs Recognize Incoherence in Narratives?
- Towards Accurate UAV Image Perception: Guiding Vision-Language Models with Stronger Task Prompts
- Social Perceptions of English Spelling Variation on Twitter: A Comparative Analysis of Human and LLM Responses
- Focused Chain-of-Thought: Efficient LLM Reasoning via Structured Input Information
- Prompt Fairness: Sub-group Disparities in LLMs
- Evaluating perturbation robustness of generative systems that use COBOL code inputs
- Cross-cultural value alignment frameworks for responsible AI governance: Evidence from China-West comparative analysis
- Can we use LLMs to bootstrap reinforcement learning? -- A case study in digital health behavior change
- On the Brittleness of LLMs: A Journey around Set Membership
- GRAPHTEXTACK: A Realistic Black-Box Node Injection Attack on LLM-Enhanced GNNs
- Thinking While Speaking: Inference-Time Knowledge Transfer for Responsive and Intelligent Conversational Voice Agents
- EvalCards: A Framework for Standardized Evaluation Reporting
- Prompting for Policy: Forecasting Macroeconomic Scenarios with Synthetic LLM Personas
- Graph-Based Alternatives to LLMs for Human Simulation
- Approximating Human Preferences Using a Multi-Judge Learned System
- Not ready for the bench: LLM legal interpretation is unstable and out of step with human judgments
- When is Routing Meaningful? Diversity and Robustness in Language Model Societies
- Will Scaling Improve Social Simulation with LLMs?
- SAGE: Stochastic Prompt Optimization via Agent-Guided Exploration
- Linear-LLM-SCM: Benchmarking LLMs for Coefficient Elicitation in Linear-Gaussian Causal Models
- Machine individuality: Separating genuine idiosyncrasy from response bias in large language models
- Thought Virus: Viral Misalignment via Subliminal Prompting in Multi-Agent Systems
- Political censorship in large language models originating from China
- State of the Art of LLM-Enabled Interaction with Visualization
- StorageXTuner: An LLM Agent-Driven Automatic Tuning Framework for Heterogeneous Storage Systems
- Compiler.next: A Search-Based Compiler to Power the AI-Native Future of Software Engineering
- ISA-Bench: Benchmarking Instruction Sensitivity for Large Audio Language Models
- When Your AI Agent Succumbs to Peer-Pressure: Studying Opinion-Change Dynamics of LLMs
- DeTAILS: Deep Thematic Analysis with Iterative LLM Support
- Mapping from Meaning: Addressing the Miscalibration of Prompt-Sensitive Language Models
- Finding Answers in Thought Matters: Revisiting Evaluation on Large Language Models with Reasoning
- Natural Language Tools: A Natural Language Approach to Tool Calling In Large Language Agents
- Flip-Flop Consistency: Unsupervised Training for Robustness to Prompt Perturbations in LLMs
- Selective Adversarial Attacks on LLM Benchmarks
- LLM Knowledge is Brittle: Truthfulness Representations Rely on Superficial Resemblance
- It's Not the Capability: Harness Sensitivity Is Non-Monotone Across LLM Agent Tiers
- If Probable, Then Acceptable? Understanding Conditional Acceptability Judgments in Large Language Models
- On Randomness in Agentic Evals
- Reward Model Perspectives: Whose Opinions Do Reward Models Reward?
- Beyond Outcome Reward: Decoupling Search and Answering Improves LLM Agents
- Activation Steering with a Feedback Controller
- Not Wrong, But Untrue: LLM Overconfidence in Document-Based Queries
- No Loss, No Gain: Gated Refinement and Adaptive Compression for Prompt Optimization
- What Does Your Benchmark Really Measure? A Framework for Robust Inference of AI Capabilities
- When Specifications Conflict: A Symmetry-Based Framework for Measuring LLM Preferences
- The Artificial Intelligence Cognitive Examination: A Survey on the Evolution of Multimodal Evaluation From Recognition to Reasoning
- Evaluating Large Language Models for Detecting Antisemitism
- How Persuasive is Your Context?
- A Taxonomy of Prompt Defects in LLM Systems
- Programmable Cognitive Bias in Social Agents
- Instance-level Randomization: Toward More Stable LLM Evaluations
- PowerChain: A Verifiable Agentic AI System for Automating Distribution Grid Analyses
- From scratch to silver: Creating trustworthy training data for patent-SDG classification using Large Language Models
- Discrimination by LLMs: Cross-lingual Bias Assessment and Mitigation in Decision-Making and Summarisation
- Agents of Discovery
- No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models
- PersonaFuse: A Personality Activation-Driven Framework for Enhancing Human-LLM Interactions
- Rethinking Reasoning in LLMs: Neuro-Symbolic Local RetoMaton Beyond ICL and CoT
- Conceptual Schema Inference for Tabular Datasets using Large Language Models
- Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases
- On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
- Knowledge Integration for Physics-informed Symbolic Regression Using Pre-trained Large Language Models
- Domain Adaptation of LLMs for Process Data
- Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs
- AHAMask: Reliable Task Specification for Large Audio Language Models without Instructions
- Compiling Prompts, Not Crafting Them: A Reproducible Workflow for AI-Assisted Evidence Synthesis
- If We May De-Presuppose: Robustly Verifying Claims through Presupposition-Free Question Decomposition
- Talk Less, Call Right: Enhancing Role-Play LLM Agents with Automatic Prompt Optimization and Role Prompting
- Inducing State Anxiety in LLM Agents Reproduces Human-Like Biases in Consumer Decision-Making
- Automatic Reviewers Fail to Detect Faulty Reasoning in Research Papers: A New Counterfactual Evaluation Framework
- Just-in-time and distributed task representations in language models
- Prompting Strategies for Language Model-Based Item Generation in K-12 Education: Bridging the Gap Between Small and Large Language Models
- Prompt Orchestration Markup Language
- Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation
- When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs
- Using Large Language Models to Measure Symptom Severity in Patients At Risk for Schizophrenia
- APIO: Automatic Prompt Induction and Optimization for Grammatical Error Correction and Text Simplification
- When Prompt Engineering Meets Software Engineering: CNL-P as Natural and Robust "APIs'' for Human-AI Interaction
- Persistent Instability in LLM's Personality Measurements: Effects of Scale, Reasoning, and Conversation History
- Pay What LLM Wants: Can LLM Simulate Economics Experiment with 522 Real-human Persona?
- Evaluating Position Bias in Large Language Model Recommendations
- Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing Styles
- Rote Learning Considered Useful: Generalizing over Memorized Data in LLMs
- Understanding Human Limits in Pattern Recognition: A Computational Model of Sequential Reasoning in Rock, Paper, Scissors
- PromptSuite: A Task-Agnostic Framework for Multi-Prompt Generation
- Agent WARPP: Workflow Adherence via Runtime Parallel Personalization
- The Levers of Political Persuasion with Conversational AI
- How LLMs Comprehend Temporal Meaning in Narratives: A Case Study in Cognitive Evaluation of LLMs
- A Unifying Scheme for Extractive Content Selection Tasks
- Grammar-Guided Evolutionary Search for Discrete Prompt Optimisation
- An Analysis of Chinese Censorship Bias in LLMs
- Anthropomimetic Uncertainty: What Verbalized Uncertainty in Language Models is Missing
- Toward Valid Measurement Of (Un)fairness For Generative AI: A Proposal For Systematization Through The Lens Of Fair Equality of Chances
- A validity-guided workflow for robust large language model research in psychology
Discussions
- Recent papers I'm reading: "How I learned to start worrying about prompt formatting" arxiv.org/abs/2310.11324 "Bridging Search and Recommendation in Generative Retrieval" dl.acm.org/doi/10.1145/... [bsky, 50 points, 2 comments]
- First, we note that the effects do not require "cognition" to be explained. The idea that they _must_ require cognition is primarily an assumption of the authors; but as seen in other research, LLM ou [bsky, 9 points, 1 comments]
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting arxiv.org/abs/2310.11324 [bsky, 4 points, 0 comments]
- LLMs are sensitive to prompt formatting too, so there's no shortage of fun confounders arxiv.org/abs/2310.11324 [bsky, 1 points, 1 comments]
Related