Large Language Models Struggle to Learn Long-Tail Knowledge
2022/11/15 by Nikhil Kandpal, Kandpal, Nikhil, Haikang Deng +7 · 2 voices · 122 citations
Computer Science · #Artificial intelligence #Code (set theory) #Computer science #Expert finding and Q&A systems #Information retrieval #Language model #Natural Language Processing Techniques #Natural language processing #Question answering #Text corpus #The Internet #Topic Modeling #Training set #World Wide Web
paper · pdf · doi:10.48550/arxiv.2211.08411
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2022/11/15 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
The Internet contains a wealth of knowledge -- from the birthdays of historical figures to tutorials on how to code -- all of which may be learned by language models. However, while certain pieces of information are ubiquitous on the web, others appear extremely rarely. In this paper, we study the relationship between the knowledge memorized by large language models and the information in pre-training datasets scraped from the web. In particular, we show that a language model's ability to answer a fact-based question relates to how many documents associated with that question were seen during pre-training. We identify these relevant documents by entity linking pre-training datasets and counting documents that contain the same entities as a given question-answer pair. Our results demonstrate strong correlational and causal relationships between accuracy and relevant document count for numerous question answering datasets (e.g., TriviaQA), pre-training corpora (e.g., ROOTS), and model sizes (e.g., 176B parameters). Moreover, while larger models are better at learning long-tail knowledge, we estimate that today's models must be scaled by many orders of magnitude to reach competitive QA performance on questions with little support in the pre-training data. Finally, we show that retrieval-augmentation can reduce the dependence on relevant pre-training information, presenting a promising approach for capturing the long-tail.
Cited by
- Reducing Hallucinations in Complex Question Answering using Simple Graph-based Retrieval-Augmented Generation (long version)
- LLM-Driven Cost-Effective Requirements Change Impact Analysis
- Grounded verification of chemical and materials reasoning: detection is the bottleneck
- Search-on-Graph: Iterative Informed Navigation for Large Language Model Reasoning on Knowledge Graphs
- Hallucinations Undermine Trust; Metacognition is a Way Forward
- Mind Your Tone: Investigating How Prompt Politeness Affects LLM Accuracy (short paper)
- How to inject knowledge efficiently? Knowledge Infusion Scaling Law for Pre-training Large Language Models
- Position: We Need An Algorithmic Understanding of Generative AI
- How Deep Do Large Language Models Internalize Scientific Literature and Citation Practices?
- The Appeal and Reality of Recycling LoRAs with Adaptive Merging
- Community size rather than grammatical complexity better predicts Large Language Model accuracy in a novel Wug Test
- Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation
- Algorithmic Blindness in Large Language Models: A Calibration Study of Performance Prediction
- DACE For Railway Acronym Disambiguation
- The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality
- Parametric Retrieval-Augmented Generation using Latent Routing of LoRA Adapters
- Data Value in the Age of Scaling: Understanding LLM Scaling Dynamics Under Real-Synthetic Data Mixtures
- BARD: budget-aware reasoning distillation
- BudgetLeak: Membership Inference Attacks on RAG Systems via the Generation Budget Side Channel
- Modeling Uncertainty Trends for Timely Retrieval in Dynamic RAG
- Hierarchical Memorization in Large Language Models: Evidence from Citation Generation
- Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention
- Place Matters: Comparing LLM Hallucination Rates for Place-Based Legal Queries
- LPFQA: A Long-Tail Professional Forum-based Benchmark for LLM Evaluation
- Reusing Pre-Training Data at Test Time is a Compute Multiplier
- Contamination Detection for VLMs using Multi-Modal Semantic Perturbation
- Kastor: Fine-tuned Small Language Models for Shape-based Active Relation Extraction
- Auditing LLM Editorial Bias in News Media Exposure
- CRAG-MM: Multi-modal Multi-turn Comprehensive RAG Benchmark
- Estimation of discrete distributions with high probability under χ2-divergence
- The Architecture of Errors: From Universal Impossibility to Patch-Local LLM Reliability
- Practical Code RAG at Scale: Task-Aware Retrieval Design Choices under Compute Budgets
- Capability Ceilings in Autoregressive Language Models: Empirical Evidence from Knowledge-Intensive Tasks
- Data-Centric Lessons To Improve Speech-Language Pretraining
- Teaming LLMs to Detect and Mitigate Hallucinations
- Facts in Stats: Impacts of Pretraining Diversity on Language Model Generalization
- Structure-R1: Dynamically Leveraging Structural Knowledge in LLM Reasoning through Reinforcement Learning
- Rewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model Behavior
- CacheClip: Accelerating RAG with Effective KV Cache Reuse
- Shaping History, Responsibly: Seven Principles to Guide the Design of Archival AI Assistants for Cultural Heritage Collections
- Abductive Preference Learning
- Large Language Models Hallucination: A Comprehensive Survey
- LongTail-Swap: benchmarking language models' abilities on rare words
- Review of Hallucination Understanding in Large Language and Vision Models
- Distributed Specialization: Rare-Token Neurons in Large Language Models
- Causal Understanding by LLMs: The Role of Uncertainty
- A mathematical theory of relational generalization in transitive inference
- Analyzing the Effects of Supervised Fine-Tuning on Model Knowledge from Token and Parameter Levels
- RelRepair: Enhancing Automated Program Repair by Retrieving Relevant Code
- Evaluating the Limitations of Local LLMs in Solving Complex Programming Challenges
- ChatGPT-generated texts show authorship traits that identify them as non-human
- No Clustering, No Routing: How Transformers Actually Process Rare Tokens
- Provable Benefits of In-Tool Learning for Large Language Models
- Predicting Failures of LLMs to Link Biomedical Ontology Terms to Identifiers Evidence Across Models and Ontologies
- LFD: Layer Fused Decoding to Exploit External Knowledge in Retrieval-Augmented Generation
- Breaking the Trade-Off Between Faithfulness and Expressiveness for Large Language Models
- Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models
- A Retrieval Augmented Spatio-Temporal Framework for Traffic Prediction
- Learning Facts at Scale with Active Reading
- Synthesizing scientific literature with retrieval-augmented language models
- Efficient Knowledge Probing of Large Language Models by Adapting Pre-trained Embeddings
- ASINT: Learning AS-to-Organization Mapping from Internet Metadata
- LinkQA: Synthesizing Diverse QA from Multiple Seeds Strongly Linked by Knowledge Points
- Fine-Grained Privacy Extraction from Retrieval-Augmented Generation Systems via Knowledge Asymmetry Exploitation
- MeMo: Memory as a Model
- LRCTI: A Large Language Model-Based Framework for Multi-Step Evidence Retrieval and Reasoning in Cyber Threat Intelligence Credibility Verification
- HedraRAG: Coordinating LLM Generation and Database Retrieval in Heterogeneous RAG Serving
- Dynamic Injection of Entity Knowledge into Dense Retrievers
- MMSearch-R1: Incentivizing LMMs to Search
- Counterfactual Influence as a Distributional Quantity
- Keeping Medical AI Healthy and Trustworthy: A Review of Detection and Correction Methods for System Degradation
- From Data to Knowledge: Evaluating How Efficiently Language Models Learn Facts
- Normative Conflicts and Shallow AI Alignment
- FictionalQA: A Dataset for Studying Memorization and Knowledge Acquisition
- Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs
- On Generalization across Measurement Systems: LLMs Entail More Test-Time Compute for Underrepresented Cultures
- Inter-Passage Verification for Multi-evidence Multi-answer QA
- OntoRAG: Enhancing Question-Answering through Automated Ontology Derivation from Unstructured Knowledge Bases
- From Parameters to Prompts: Understanding and Mitigating the Factuality Gap between Fine-Tuned LLMs
- Beyond Monoliths: Expert Orchestration for More Capable, Democratic, and Safe Language Models
- Prompting is not Enough: Exploring Knowledge Integration and Controllable Generation
- MA-RAG: Multi-Agent Retrieval-Augmented Generation via Collaborative Chain-of-Thought Reasoning
- Knoll: Creating a Knowledge Ecosystem for Large Language Models
- Language Model Behavior: A Comprehensive Survey
- BiomedSQL: Text-to-SQL for Scientific Reasoning on Biomedical Knowledge Bases
- Data Mixing Can Induce Phase Transitions in Knowledge Acquisition
- Distilling LLM Agent into Small Models with Retrieval and Code Tools
- Diagnosing our datasets: How does my language model learn clinical information?
- Enhancing LLMs via High-Knowledge Data Selection
- Divide by Question, Conquer by Agent: SPLIT-RAG with Question-Driven Graph Partitioning
- Emergent Specialization: Rare Token Neurons in Language Models
- GAP: Graph-Assisted Prompts for Dialogue-based Medication Recommendation
- From n-gram to Attention: How Model Architectures Learn and Propagate Bias in Language Modeling
- Learning Auxiliary Tasks Improves Reference-Free Hallucination Detection in Open-Domain Long-Form Generation
- DACL-RAG: Data Augmentation Strategy with Curriculum Learning for Retrieval-Augmented Generation
- IterKey: Iterative Keyword Generation with LLMs for Enhanced Retrieval Augmented Generation
- Bridging AI and Carbon Capture: A Dataset for LLMs in Ionic Liquids and CBE Research
- CHORUS: Zero-shot Hierarchical Retrieval and Orchestration for Generating Linear Programming Code
- Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency
- The Model Knows Your Project, Not You: Measuring Recognition in LLMs with NameRank
- A Glitch in the Matrix? Locating and Detecting Language Model Grounding with Fakepedia
- Natural Ungrokking: Asymmetric Control of Which Rules Survive Pretraining
- EnronQA: Towards Personalized RAG over Private Documents
- The Synthetic Web: Adversarially-Curated Mini-Internets for Diagnosing Epistemic Weaknesses of Language Agents
- Evolutionary Context Search for Automated Skill Acquisition
- Long-Tail Knowledge in Large Language Models: Taxonomy, Mechanisms, Interventions and Implications
- MTR-Suite: A Framework for Evaluating and Synthesizing Conversational Retrieval Benchmarks
- Co-LMLM: Continuous-Query Limited Memory Language Models
- Can LLMs Introspect? A Reality Check
- NanoKnow: How to Know What Your Language Model Knows
- Enhancing Privacy-Utility Trade-offs to Mitigate Memorization in Diffusion Models
- HalluLens: LLM Hallucination Benchmark
- T1: Tool-integrated Verification for Test-time Compute Scaling in Small Language Models
- FinDER: Financial Dataset for Question Answering and Evaluating Retrieval-Augmented Generation
- CoLoTa: A Dataset for Entity-based Commonsense Reasoning over Long-Tail Knowledge
- Hypothetical Documents or Knowledge Leakage? Rethinking LLM-based Query Expansion
- FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents
- The Other Side of the Coin: Exploring Fairness in Retrieval-Augmented Generation
- RiskRAG: A Data-Driven Solution for Improved AI Model Risk Reporting
- Efficient Tuning of Large Language Models for Knowledge-Grounded Dialogue Generation
- A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility
- Retrieval Augmented Generation with Collaborative Filtering for Personalized Text Generation
Discussions
Related