When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories
2022/12/20 by Alex Mallen, Mallen, Alex, Akari Asai +9 · 203 citations
Computer Science · Decision Sciences · #Topic Modeling #Natural Language Processing Techniques #Data Quality and Management
paper · pdf · doi:10.48550/arxiv.2212.10511
Abstract
Despite their impressive performance on diverse tasks, large language models (LMs) still struggle with tasks requiring rich world knowledge, implying the limitations of relying solely on their parameters to encode a wealth of world knowledge. This paper aims to understand LMs' strengths and limitations in memorizing factual knowledge, by conducting large-scale knowledge probing experiments of 10 models and 4 augmentation methods on PopQA, our new open-domain QA dataset with 14k questions. We find that LMs struggle with less popular factual knowledge, and that scaling fails to appreciably improve memorization of factual knowledge in the long tail. We then show that retrieval-augmented LMs largely outperform orders of magnitude larger LMs, while unassisted LMs remain competitive in questions about high-popularity entities. Based on those findings, we devise a simple, yet effective, method for powerful and efficient retrieval-augmented LMs, which retrieves non-parametric memories only when necessary. Experimental results show that this significantly improves models' performance while reducing the inference costs.
Cited by
- From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search
- FoldAct: Efficient and Stable Context Folding for Long-Horizon Search Agents
- EviBack: Search-Agent Reinforcement Learning via Evidence-Constrained Teacher Backoff
- Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation
- TRACE: Business Rule-Grounded Reasoning Curriculum for Knowledge-Preserving Parametric Tool Retrieval in Enterprise LLMs
- Exploring the Security Threats of Retriever Backdoors in Retrieval-Augmented Code Generation
- QuCo-RAG: Quantifying Uncertainty from the Pre-training Corpus for Dynamic Retrieval-Augmented Generation
- KV Admission: Learning What to Write for Efficient Long-Context Inference
- AdaSearch: Balancing Parametric Knowledge and Search in Large Language Models via Reinforcement Learning
- PDE-Agent: A toolchain-augmented multi-agent framework for PDE solving
- ToolForge: A Data Synthesis Pipeline for Multi-Hop Search without Real-World APIs
- Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
- CoDA: A Context-Decoupled Hierarchical Agent with Reinforcement Learning
- MedAI: Evaluating TxAgent's Therapeutic Agentic Reasoning in the NeurIPS CURE-Bench Competition
- PathFinder: MCTS and LLM Feedback-based Path Selection for Multi-Hop Question Answering
- RouteRAG: Efficient Retrieval-Augmented Generation from Text and Graph via Reinforcement Learning
- Enhancing Agentic RL with Progressive Reward Shaping and Value-based Sampling Policy Optimization
- Faithfulness metric fusion: Improving the evaluation of LLM trustworthiness across domains
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
- CARL: Criticality-Aware Agentic Reinforcement Learning
- Spatially-Enhanced Retrieval-Augmented Generation for Walkability and Urban Discovery
- On GRPO Collapse in Search-R1: The Lazy Likelihood-Displacement Death Spiral
- BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents
- Towards Unification of Hallucination Detection and Fact Verification for Large Language Models
- Agentic Policy Optimization via Instruction-Policy Co-Evolution
- Training Introspective Behavior: Fine-Tuning Induces Reliable Internal State Detection in a 7B Model
- Context-Aware Pragmatic Metacognitive Prompting for Sarcasm Detection
- TrackList: Tracing Back Query Linguistic Diversity for Head and Tail Knowledge in Open Large Language Models
- Concept than Document: Context Compression via AMR-based Conceptual Entropy
- HyperbolicRAG: Enhancing Retrieval-Augmented Generation with Hyperbolic Representations
- Parametric Retrieval-Augmented Generation using Latent Routing of LoRA Adapters
- Global Resolution: Optimal Multi-Draft Speculative Sampling via Convex Minimization
- Noise-Robust Abstractive Compression in Retrieval-Augmented Language Models
- LiveRAG: A diverse Q&A dataset with varying difficulty level for RAG evaluation
- NeuroPath: Neurobiology-Inspired Path Tracking and Reflection for Semantically Coherent Retrieval
- HEDGE: Hallucination Estimation via Dense Geometric Entropy for VQA with Vision-Language Models
- Retrieval as a Decision: Training-Free Adaptive Gating for Efficient RAG
- Thinking Forward and Backward: Multi-Objective Reinforcement Learning for Retrieval-Augmented Reasoning
- Learning to Focus: Focal Attention for Selective and Scalable Transformers
- TabRAG: Tabular Document Retrieval via Structured Language Representations
- LPFQA: A Long-Tail Professional Forum-based Benchmark for LLM Evaluation
- Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs
- The Illusion of Certainty: Uncertainty Quantification for LLMs Fails under Ambiguity
- Interpreting Multi-Attribute Confounding through Numerical Attributes in Large Language Models
- MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning
- Rescuing the Unpoisoned: Efficient Defense against Knowledge Corruption Attacks on RAG Systems
- Efficient Test-Time Retrieval Augmented Generation
- Prompt-R1: Collaborative Automatic Prompting Framework via End-to-end Reinforcement Learning
- Separate the Wheat from the Chaff: Winnowing Down Divergent Views in Retrieval Augmented Generation
- Interact-RAG: Reason and Interact with the Corpus, Beyond Black-Box Retrieval
- InfoFlow: Reinforcing Search Agent Via Reward Density Optimization
- Graph-Enhanced Policy Optimization in LLM Agent Training
- Towards Global Retrieval Augmented Generation: A Benchmark for Corpus-Level Reasoning
- CRAG-MM: Multi-modal Multi-turn Comprehensive RAG Benchmark
- GAP: Graph-Based Agent Planning with Parallel Tool Use and Reinforcement Learning
- Beyond Self-Knowledge: Propagating Uncertainty Across Reasoning and Retrieval in LLMs
- SimpleWikiSearch: A Clean Offline Wikipedia Environment for Agentic Search
- The Architecture of Errors: From Universal Impossibility to Patch-Local LLM Reliability
- Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
- Sharpness-Guided Group Relative Policy Optimization via Probability Shaping
- Repurposing Synthetic Data for Fine-grained Search Agent Supervision
- Optimizing Retrieval for RAG via Reinforced Contrastive Learning
- MAD-Fact: A Multi-Agent Debate Framework for Long-Form Factuality Evaluation in LLMs
- RaCoT: Plug-and-Play Contrastive Example Generation Mechanism for Enhanced LLM Reasoning Reliability
- Redefining Retrieval Evaluation in the Era of LLMs
- Bridging Language Gaps with Adaptive RAG: Improving Indonesian Language Question Answering
- Neural Diversity Regularizes Hallucinations in Small Models
- From Facts to Folklore: Evaluating Large Language Models on Bengali Cultural Knowledge
- When Facts Change: Probing LLMs on Evolving Knowledge with evolveQA
- See the Text: From Tokenization to Visual Reading
- Search Self-play: Pushing the Frontier of Agent Capability without Supervision
- WebSeer: Training Deeper Search Agents through Reinforcement Learning with Self-Reflection
- ChronoPlay: A Framework for Modeling Dual Dynamics and Authenticity in Game RAG Benchmarks
- Annotation-Efficient Universal Honesty Alignment
- A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Evaluations, and Applications
- AutoGraph-R1: End-to-End Reinforcement Learning for Knowledge Graph Construction
- Facts in Stats: Impacts of Pretraining Diversity on Language Model Generalization
- Cost-Aware Retrieval-Augmentation Reasoning Models with Adaptive Retrieval Depth
- EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle
- Structure-R1: Dynamically Leveraging Structural Knowledge in LLM Reasoning through Reinforcement Learning
- Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn LLM Agents
- Stop-RAG: Value-Based Retrieval Control for Iterative RAG
- Rewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model Behavior
- Towards Agentic Self-Learning LLMs in Search Environment
- BRIEF-Pro: Universal Context Compression with Short-to-Long Synthesis for Fast and Accurate Multi-Hop Reasoning
- Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation
- MedREK: Retrieval-Based Editing for Medical LLMs with Key-Aware Prompts
- LLM-guided Hierarchical Search for End-to-end Reasoning Intensive Retrieval
- Understanding Parametric Knowledge Injection in Retrieval-Augmented Generation
- Teaching Language Models to Faithfully Express their Uncertainty
- Uncertainty Quantification for Retrieval-Augmented Reasoning
- Domain-Specific Data Generation Framework for RAG Adaptation
- BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions
- RECON: Reasoning with Condensation for Efficient Retrieval-Augmented Generation
- Beyond the limitation of a single query: Train your LLM for query expansion with Reinforcement Learning
- Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipe
- Don't Throw Away Your Pretrained Model
- DSPO: Stable and Efficient Policy Optimization for Agentic Search and Reasoning
- When Retrieval Succeeds and Fails: Rethinking Retrieval-Augmented Generation for LLMs
- Large Language Models Do NOT Really Know What They Don't Know
- HiPRAG: Hierarchical Process Rewards for Efficient Agentic Retrieval Augmented Generation
- SUBQRAG: Sub-Question Driven Dynamic Graph RAG
- A2Search: Ambiguity-Aware Question Answering with Reinforcement Learning
- A Comparison of Independent and Joint Fine-tuning Strategies for Retrieval-Augmented Generation
- Beneficial Reasoning Behaviors in Agentic Search and Effective Post-training to Obtain Them
- Stratified GRPO: Handling Structural Heterogeneity in Reinforcement Learning of LLM Search Agents
- DecEx-RAG: Boosting Agentic Retrieval-Augmented Generation with Decision and Execution Optimization via Process Supervision
- MARS: Co-evolving Dual-System Deep Research via Multi-Agent Reinforcement Learning
- Beyond Outcome Reward: Decoupling Search and Answering Improves LLM Agents
- MHA-RAG: Improving Efficiency, Accuracy, and Consistency by Encoding Exemplars as Soft Prompts
- Equipping Retrieval-Augmented Large Language Models with Document Structure Awareness
- Sample, Align, Synthesize: Graph-Based Response Synthesis with ConGrs
- HalluGuard: Evidence-Grounded Small Reasoning Models to Mitigate Hallucinations in Retrieval-Augmented Generation
- ReSeek: A Self-Correcting Framework for Search Agents with Instructive Rewards
- RE-Searcher: Robust Agentic Search with Goal-oriented Planning and Self-reflection
- Hybrid Reward Normalization for Process-supervised Non-verifiable Agentic Tasks
- FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology
- Beyond Static Retrieval: Opportunities and Pitfalls of Iterative Retrieval in GraphRAG
- MemGen: Weaving Generative Latent Memory for Self-Evolving Agents
- Training Dynamics of Parametric and In-Context Knowledge Utilization in Language Models
- Can Large Language Models Express Uncertainty Like Human?
- AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play
- Knowledge Homophily in Large Language Models
- Retrieval-Constrained Decoding Reveals Underestimated Parametric Knowledge in Language Models
- Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning
- Do LLM Agents Know How to Ground, Recover, and Assess? A Benchmark for Epistemic Competence in Information-Seeking Agents
- Think Right, Not More: Test-Time Scaling for Numerical Claim Verification
- Black-Box Hallucination Detection via Consistency Under the Uncertain Expression
- Hallucination reduction with CASAL: Contrastive Activation Steering For Amortized Learning
- Tree Search for LLM Agent Reinforcement Learning
- Harness-G: A Graph-Structured Harness for Search Agents
- Group-Reflective Self-Distillation for Agentic Reinforcement Learning
- Models for minimalist RAG: B1ade 335M Embedding and 1B Parameter Small Language Models
- CoCoA: Confidence and Context-Aware Adaptive Decoding for Resolving Knowledge Conflicts in Large Language Models
- AttnComp: Attention-Guided Adaptive Context Compression for Retrieval-Augmented Generation
- SignalLLM: A General-Purpose LLM Agent Framework for Automated Signal Processing
- Quantifying Self-Awareness of Knowledge in Large Language Models
- Thinking in a Crowd: How Auxiliary Information Shapes LLM Reasoning
- DSCC-HS: A Dynamic Self-Reinforcing Framework for Hallucination Suppression in Large Language Models
- Improving Context Fidelity via Native Retrieval-Augmented Reasoning
- LLMs Can't Handle Peer Pressure: Crumbling under Multi-Agent Social Interactions
- InfoGain-RAG: Boosting Retrieval-Augmented Generation via Document Information Gain-based Reranking and Filtering
- Harnessing Optimization Dynamics for Curvature-Informed Model Merging
- Explaining Black-box Language Models with Knowledge Probing Systems: A Post-hoc Explanation Perspective
- Unbiased Reasoning for Knowledge-Intensive Tasks in Large Language Models via Conditional Front-Door Adjustment
- AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
- HANRAG: Heuristic Accurate Noise-resistant Retrieval-Augmented Generation for Multi-hop Question Answering
- Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
- MeVe: A Modular System for Memory Verification and Effective Context Control in Language Models
- Privacy-Preserving Reasoning with Knowledge-Distilled Parametric Retrieval Augmented Generation
- VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
- EviNote-RAG: Enhancing RAG Models via Answer-Supportive Evidence Notes
- Open Data Synthesis For Deep Research
- AI-SearchPlanner: Modular Agentic Search via Pareto-Optimal Multi-Objective Reinforcement Learning
- Can Compact Language Models Search Like Agents? Distillation-Guided Policy Optimization for Preserving Agentic RAG Capabilities
- Continuously Steering LLMs Sensitivity to Contextual Knowledge with Proxy Models
- Test-time Corpus Feedback: From Retrieval to RAG
- LongRecall: A Structured Approach for Robust Recall Evaluation in Long-Form Text
- TracSum: A New Benchmark for Aspect-Based Summarization with Sentence-Level Traceability in Medical Domain
- CardAIc-Agents: A Multimodal Framework with Hierarchical Adaptation for Cardiac Care Support
- Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward
- DynamixSFT: Dynamic Mixture Optimization of Instruction Tuning Collections
- DiFaR: Enhancing Multimodal Misinformation Detection with Diverse, Factual, and Relevant Rationales
- Towards Self-cognitive Exploration: Metacognitive Knowledge Graph Retrieval Augmented Generation
- ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning
- Can LLMs Detect Their Confabulations? Estimating Reliability in Uncertainty-Aware Language Models
- Improving Document Retrieval Coherence for Semantically Equivalent Queries
- Beyond Ten Turns: Unlocking Long-Horizon Agentic Search with Large-Scale Asynchronous RL
- Synthesizing scientific literature with retrieval-augmented language models
- Impact-driven Context Filtering For Cross-file Code Completion
- VeriGUI: Verifiable Long-Chain GUI Dataset
- An Entity Linking Agent for Question Answering
- CF-RAG: A Dataset and Method for Carbon Footprint QA Using Retrieval-Augmented Generation
- Key-Augmented Neural Triggers for Knowledge Sharing
- Token-Level Precise Attack on RAG: Searching for the Best Alternatives to Mislead Generation
- Simple Methods Defend RAG Systems Well Against Real-World Attacks
- Beyond Chunks and Graphs: Retrieval-Augmented Generation through Triplet-Driven Thinking
- Harnessing Temporal Databases for Systematic Evaluation of Factual Time-Sensitive Question-Answering in Large Language Models
- LMAR: Language Model Augmented Retriever for Domain-specific Knowledge Indexing
- CoCoA: Collaborative Chain-of-Agents for Parametric-Retrieved Knowledge Synergy
- A Survey of LLM-based Deep Search Agents: Paradigm, Optimization, Evaluation, and Challenges
- MAO-ARAG: Multi-Agent Orchestration for Adaptive Retrieval-Augmented Generation
- DiffLoRA: Differential Low-Rank Adapters for Large Language Models
- GraphRAG-R1: Graph Retrieval-Augmented Generation with Process-Constrained Reinforcement Learning
- Graph-R1: Towards Agentic GraphRAG Framework via End-to-end Reinforcement Learning
- RAG in the Wild: On the (In)effectiveness of LLMs with Mixture-of-Knowledge Retrieval Augmentation
- Injecting External Knowledge into the Reasoning Process Enhances Retrieval-Augmented Generation
- Deceptive Grounding: Entity Attribution Failure in Clinical Retrieval-Augmented Generation
- Mitigating Factual Hallucination in Large Reasoning Models via Mixed-Mode Advantage Regularization
- PhantomBench: Benchmarking the Non-existential Threat of Language Models
- ERNIE 5.0 Technical Report
- PRGB Benchmark: A Robust Placeholder-Assisted Algorithm for Benchmarking Retrieval-Augmented Generation
- Aligning Knowledge Graphs and Language Models for Factual Accuracy
- The Curious Case of Factuality Finetuning: Models' Internal Beliefs Can Improve Factuality
- Clue-RAG: Towards Accurate and Cost-Efficient Graph-based RAG via Multi-Partite Graph and Query-Driven Iterative Retrieval
- Shifting from Ranking to Set Selection for Retrieval Augmented Generation
- The Delta Learning Hypothesis: Preference Tuning on Weak Data can Yield Strong Gains
- HIRAG: Hierarchical-Thought Instruction-Tuning Retrieval-Augmented Generation
- Dynamic Injection of Entity Knowledge into Dense Retrievers
- Thinking About Thinking: SAGE-nano's Inverse Reasoning for Self-Aware Language Models
- RAG-R1: Incentivizing the Search and Reasoning Capabilities of LLMs through Multi-query Parallelism
- Response Quality Assessment for Retrieval-Augmented Generation via Conditional Conformal Factuality
- EraRAG: Efficient and Incremental Retrieval Augmented Generation for Growing Corpora
Related