When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories
2022/12/20 by Alex Mallen, Mallen, Alex, Akari Asai +9 · 371 citations
Computer Science · Decision Sciences · #Topic Modeling #Natural Language Processing Techniques #Data Quality and Management
paper · pdf · doi:10.48550/arxiv.2212.10511
Abstract
Despite their impressive performance on diverse tasks, large language models (LMs) still struggle with tasks requiring rich world knowledge, implying the limitations of relying solely on their parameters to encode a wealth of world knowledge. This paper aims to understand LMs' strengths and limitations in memorizing factual knowledge, by conducting large-scale knowledge probing experiments of 10 models and 4 augmentation methods on PopQA, our new open-domain QA dataset with 14k questions. We find that LMs struggle with less popular factual knowledge, and that scaling fails to appreciably improve memorization of factual knowledge in the long tail. We then show that retrieval-augmented LMs largely outperform orders of magnitude larger LMs, while unassisted LMs remain competitive in questions about high-popularity entities. Based on those findings, we devise a simple, yet effective, method for powerful and efficient retrieval-augmented LMs, which retrieves non-parametric memories only when necessary. Experimental results show that this significantly improves models' performance while reducing the inference costs.
Cited by
- From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search
- FoldAct: Efficient and Stable Context Folding for Long-Horizon Search Agents
- EviBack: Search-Agent Reinforcement Learning via Evidence-Constrained Teacher Backoff
- Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation
- TRACE: Business Rule-Grounded Reasoning Curriculum for Knowledge-Preserving Parametric Tool Retrieval in Enterprise LLMs
- Exploring the Security Threats of Retriever Backdoors in Retrieval-Augmented Code Generation
- QuCo-RAG: Quantifying Uncertainty from the Pre-training Corpus for Dynamic Retrieval-Augmented Generation
- KV Admission: Learning What to Write for Efficient Long-Context Inference
- AdaSearch: Balancing Parametric Knowledge and Search in Large Language Models via Reinforcement Learning
- PDE-Agent: A toolchain-augmented multi-agent framework for PDE solving
- ToolForge: A Data Synthesis Pipeline for Multi-Hop Search without Real-World APIs
- Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
- CoDA: A Context-Decoupled Hierarchical Agent with Reinforcement Learning
- MedAI: Evaluating TxAgent's Therapeutic Agentic Reasoning in the NeurIPS CURE-Bench Competition
- PathFinder: MCTS and LLM Feedback-based Path Selection for Multi-Hop Question Answering
- RouteRAG: Efficient Retrieval-Augmented Generation from Text and Graph via Reinforcement Learning
- Enhancing Agentic RL with Progressive Reward Shaping and Value-based Sampling Policy Optimization
- Faithfulness metric fusion: Improving the evaluation of LLM trustworthiness across domains
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
- CARL: Criticality-Aware Agentic Reinforcement Learning
- Spatially-Enhanced Retrieval-Augmented Generation for Walkability and Urban Discovery
- On Group Relative Policy Optimization Collapse in Agent Search: The Lazy Likelihood-Displacement
- BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents
- Towards Unification of Hallucination Detection and Fact Verification for Large Language Models
- Agentic Policy Optimization via Instruction-Policy Co-Evolution
- Steering Awareness: Detecting Activation Steering from Within
- Context-Aware Pragmatic Metacognitive Prompting for Sarcasm Detection
- TrackList: Tracing Back Query Linguistic Diversity for Head and Tail Knowledge in Open Large Language Models
- Concept than Document: Context Compression via AMR-based Conceptual Entropy
- HyperbolicRAG: Enhancing Retrieval-Augmented Generation with Hyperbolic Representations
- Parametric Retrieval-Augmented Generation using Latent Routing of LoRA Adapters
- Global Resolution: Optimal Multi-Draft Speculative Sampling via Convex Minimization
- Noise-Robust Abstractive Compression in Retrieval-Augmented Language Models
- LiveRAG: A diverse Q&A dataset with varying difficulty level for RAG evaluation
- NeuroPath: Neurobiology-Inspired Path Tracking and Reflection for Semantically Coherent Retrieval
- HEDGE: Hallucination Estimation via Dense Geometric Entropy for VQA with Vision-Language Models
- Retrieval as a Decision: Training-Free Adaptive Gating for Efficient RAG
- Thinking Forward and Backward: Multi-Objective Reinforcement Learning for Retrieval-Augmented Reasoning
- Learning to Focus: Focal Attention for Selective and Scalable Transformers
- TabRAG: Improving Tabular Document Question Answering for Retrieval Augmented Generation via Structured Representations
- LPFQA: A Long-Tail Professional Forum-based Benchmark for LLM Evaluation
- Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs
- The Illusion of Certainty: Uncertainty Quantification for LLMs Fails under Ambiguity
- Interpreting Multi-Attribute Confounding through Numerical Attributes in Large Language Models
- MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning
- Rescuing the Unpoisoned: Efficient Defense against Knowledge Corruption Attacks on RAG Systems
- Efficient Test-Time Retrieval Augmented Generation
- Prompt-R1: Collaborative Automatic Prompting Framework via End-to-end Reinforcement Learning
- Separate the Wheat from the Chaff: Winnowing Down Divergent Views in Retrieval Augmented Generation
- Interact-RAG: Reason and Interact with the Corpus, Beyond Black-Box Retrieval
- InfoFlow: Reinforcing Search Agent Via Reward Density Optimization
- Graph-Enhanced Policy Optimization in LLM Agent Training
- Towards Global Retrieval Augmented Generation: A Benchmark for Corpus-Level Reasoning
- CRAG-MM: Multi-modal Multi-turn Comprehensive RAG Benchmark
- GAP: Graph-Based Agent Planning with Parallel Tool Use and Reinforcement Learning
- Beyond Self-Knowledge: Propagating Uncertainty Across Reasoning and Retrieval in LLMs
- SimpleWikiSearch: A Clean Offline Wikipedia Environment for Agentic Search
- The Architecture of Errors: From Universal Impossibility to Patch-Local LLM Reliability
- Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
- Sharpness-Guided Group Relative Policy Optimization via Probability Shaping
- Repurposing Synthetic Data for Fine-grained Search Agent Supervision
- Optimizing Retrieval for RAG via Reinforced Contrastive Learning
- MAD-Fact: A Multi-Agent Debate Framework for Long-Form Factuality Evaluation in LLMs
- RaCoT: Plug-and-Play Contrastive Example Generation Mechanism for Enhanced LLM Reasoning Reliability
- Redefining Retrieval Evaluation in the Era of LLMs
- Bridging Language Gaps with Adaptive RAG: Improving Indonesian Language Question Answering
- Neural Diversity Regularizes Hallucinations in Language Models
- From Facts to Folklore: Evaluating Large Language Models on Bengali Cultural Knowledge
- When Facts Change: Probing LLMs on Evolving Knowledge with evolveQA
- See the Text: From Tokenization to Visual Reading
- Search Self-play: Pushing the Frontier of Agent Capability without Supervision
- WebSeer: Training Deeper Search Agents through Reinforcement Learning with Self-Reflection
- ChronoPlay: A Framework for Modeling Dual Dynamics and Authenticity in Game RAG Benchmarks
- Annotation-Efficient Universal Honesty Alignment
- A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Evaluations, and Applications
- AutoGraph-R1: End-to-End Reinforcement Learning for Knowledge Graph Construction
- Facts in Stats: Impacts of Pretraining Diversity on Language Model Generalization
- Cost-Aware Retrieval-Augmentation Reasoning Models with Adaptive Retrieval Depth
- EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle
- Structure-R1: Dynamically Leveraging Structural Knowledge in LLM Reasoning through Reinforcement Learning
- Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search Agents
- Stop-RAG: Value-Based Retrieval Control for Iterative RAG
- Rewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model Behavior
- Towards Agentic Self-Learning LLMs in Search Environment
- BRIEF-Pro: Universal Context Compression with Short-to-Long Synthesis for Fast and Accurate Multi-Hop Reasoning
- Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation
- MedREK: Retrieval-Based Editing for Medical LLMs with Key-Aware Prompts
- LLM-guided Hierarchical Search for End-to-end Reasoning Intensive Retrieval
- Understanding Parametric Knowledge Injection in Retrieval-Augmented Generation
- Teaching Language Models to Faithfully Express their Uncertainty
- Uncertainty Quantification for Retrieval-Augmented Reasoning
- Domain-Specific Data Generation Framework for RAG Adaptation
- BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions
- RECON: Reasoning with Condensation for Efficient Retrieval-Augmented Generation
- Beyond the limitation of a single query: Train your LLM for query expansion with Reinforcement Learning
- Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipe
- Don't Throw Away Your Pretrained Model
- DSPO: Stable and Efficient Policy Optimization for Agentic Search and Reasoning
- When Retrieval Succeeds and Fails: Rethinking Retrieval-Augmented Generation for LLMs
- Large Language Models Do NOT Really Know What They Don't Know
- HiPRAG: Hierarchical Process Rewards for Efficient Agentic Retrieval Augmented Generation
- SUBQRAG: Sub-Question Driven Dynamic Graph RAG
- A2Search: Ambiguity-Aware Question Answering with Reinforcement Learning
- A Comparison of Independent and Joint Fine-tuning Strategies for Retrieval-Augmented Generation
- Beneficial Reasoning Behaviors in Agentic Search and Effective Post-training to Obtain Them
- Stratified GRPO: Handling Structural Heterogeneity in Reinforcement Learning of LLM Search Agents
- DecEx-RAG: Boosting Agentic Retrieval-Augmented Generation with Decision and Execution Optimization via Process Supervision
- MARS: Co-evolving Dual-System Deep Research via Multi-Agent Reinforcement Learning
- Beyond Outcome Reward: Decoupling Search and Answering Improves LLM Agents
- MHA-RAG: Improving Efficiency, Accuracy, and Consistency by Encoding Exemplars as Soft Prompts
- Equipping Retrieval-Augmented Large Language Models with Document Structure Awareness
- Sample, Align, Synthesize: Graph-Based Response Synthesis with ConGrs
- HalluGuard: Evidence-Grounded Small Reasoning Models to Mitigate Hallucinations in Retrieval-Augmented Generation
- ReSeek: A Self-Correcting Framework for Search Agents with Instructive Rewards
- RE-Searcher: Robust Agentic Search with Goal-oriented Planning and Self-reflection
- Hybrid Reward Normalization for Process-supervised Non-verifiable Agentic Tasks
- FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology
- Beyond Static Retrieval: Opportunities and Pitfalls of Iterative Retrieval in GraphRAG
- MemGen: Weaving Generative Latent Memory for Self-Evolving Agents
- Training Dynamics of Parametric and In-Context Knowledge Utilization in Language Models
- Can Large Language Models Express Uncertainty Like Human?
- AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play
- Knowledge Homophily in Large Language Models
- Retrieval-Constrained Decoding Reveals Underestimated Parametric Knowledge in Language Models
- Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning
- Do LLM Agents Know How to Ground, Recover, and Assess? A Benchmark for Epistemic Competence in Information-Seeking Agents
- Think Right, Not More: Test-Time Scaling for Numerical Claim Verification
- Black-Box Hallucination Detection via Consistency Under the Uncertain Expression
- Hallucination reduction with CASAL: Contrastive Activation Steering For Amortized Learning
- Tree Search for LLM Agent Reinforcement Learning
- Harness-G: A Graph-Structured Harness for Search Agents
- Group-Reflective Self-Distillation for Agentic Reinforcement Learning
- Models for minimalist RAG: B1ade 335M Embedding and 1B Parameter Small Language Models
- CoCoA: Confidence and Context-Aware Adaptive Decoding for Resolving Knowledge Conflicts in Large Language Models
- AttnComp: Attention-Guided Adaptive Context Compression for Retrieval-Augmented Generation
- SignalLLM: A General-Purpose LLM Agent Framework for Automated Signal Processing
- Quantifying Self-Awareness of Knowledge in Large Language Models
- Thinking in a Crowd: How Auxiliary Information Shapes LLM Reasoning
- DSCC-HS: A Dynamic Self-Reinforcing Framework for Hallucination Suppression in Large Language Models
- Improving Context Fidelity via Native Retrieval-Augmented Reasoning
- LLMs Can't Handle Peer Pressure: Crumbling under Multi-Agent Social Interactions
- InfoGain-RAG: Boosting Retrieval-Augmented Generation via Document Information Gain-based Reranking and Filtering
- Harnessing Optimization Dynamics for Curvature-Informed Model Merging
- Explaining Black-box Language Models with Knowledge Probing Systems: A Post-hoc Explanation Perspective
- Unbiased Reasoning for Knowledge-Intensive Tasks in Large Language Models via Conditional Front-Door Adjustment
- AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
- HANRAG: Heuristic Accurate Noise-resistant Retrieval-Augmented Generation for Multi-hop Question Answering
- Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
- MeVe: A Modular System for Memory Verification and Effective Context Control in Language Models
- Privacy-Preserving Reasoning with Knowledge-Distilled Parametric Retrieval Augmented Generation
- VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
- EviNote-RAG: Enhancing RAG Models via Answer-Supportive Evidence Notes
- Open Data Synthesis For Deep Research
- AI-SearchPlanner: Modular Agentic Search via Pareto-Optimal Multi-Objective Reinforcement Learning
- Can Compact Language Models Search Like Agents? Distillation-Guided Policy Optimization for Preserving Agentic RAG Capabilities
- Continuously Steering LLMs Sensitivity to Contextual Knowledge with Proxy Models
- Test-time Corpus Feedback: From Retrieval to RAG
- LongRecall: A Structured Approach for Robust Recall Evaluation in Long-Form Text
- TracSum: A New Benchmark for Aspect-Based Summarization with Sentence-Level Traceability in Medical Domain
- CardAIc-Agents: A Multimodal Framework with Hierarchical Adaptation for Cardiac Care Support
- Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward
- DynamixSFT: Dynamic Mixture Optimization of Instruction Tuning Collections
- DiFaR: Enhancing Multimodal Misinformation Detection with Diverse, Factual, and Relevant Rationales
- Towards Self-cognitive Exploration: Metacognitive Knowledge Graph Retrieval Augmented Generation
- ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning
- Can LLMs Detect Their Confabulations? Estimating Reliability in Uncertainty-Aware Language Models
- Improving Document Retrieval Coherence for Semantically Equivalent Queries
- Beyond Ten Turns: Unlocking Long-Horizon Agentic Search with Large-Scale Asynchronous RL
- Synthesizing scientific literature with retrieval-augmented language models
- Impact-driven Context Filtering For Cross-file Code Completion
- VeriGUI: Verifiable Long-Chain GUI Dataset
- An Entity Linking Agent for Question Answering
- CF-RAG: A Dataset and Method for Carbon Footprint QA Using Retrieval-Augmented Generation
- Key-Augmented Neural Triggers for Knowledge Sharing
- Token-Level Precise Attack on RAG: Searching for the Best Alternatives to Mislead Generation
- Simple Methods Defend RAG Systems Well Against Real-World Attacks
- Beyond Chunks and Graphs: Retrieval-Augmented Generation through Triplet-Driven Thinking
- Harnessing Temporal Databases for Systematic Evaluation of Factual Time-Sensitive Question-Answering in Large Language Models
- LMAR: Language Model Augmented Retriever for Domain-specific Knowledge Indexing
- CoCoA: Collaborative Chain-of-Agents for Parametric-Retrieved Knowledge Synergy
- A Survey of LLM-based Deep Search Agents: Paradigm, Optimization, Evaluation, and Challenges
- MAO-ARAG: Multi-Agent Orchestration for Adaptive Retrieval-Augmented Generation
- DiffLoRA: Differential Low-Rank Adapters for Large Language Models
- GraphRAG-R1: Graph Retrieval-Augmented Generation with Process-Constrained Reinforcement Learning
- Graph-R1: Towards Agentic GraphRAG Framework via End-to-end Reinforcement Learning
- RAG in the Wild: On the (In)effectiveness of LLMs with Mixture-of-Knowledge Retrieval Augmentation
- Injecting External Knowledge into the Reasoning Process Enhances Retrieval-Augmented Generation
- Deceptive Grounding: Entity Attribution Failure in Clinical Retrieval-Augmented Generation
- Mitigating Factual Hallucination in Large Reasoning Models via Mixed-Mode Advantage Regularization
- PhantomBench: Benchmarking the Non-existential Threat of Language Models
- ERNIE 5.0 Technical Report
- PRGB Benchmark: A Robust Placeholder-Assisted Algorithm for Benchmarking Retrieval-Augmented Generation
- Aligning Knowledge Graphs and Language Models for Factual Accuracy
- Router-R1: Teaching LLMs Multi-Round Routing and Aggregation via Reinforcement Learning
- Reinforcement Fine-Tuning for Reasoning towards Multi-Step Multi-Source Search in Large Language Models
- DRAGged into Conflicts: Detecting and Addressing Conflicting Sources in Search-Augmented LLMs
- The Curious Case of Factuality Finetuning: Models' Internal Beliefs Can Improve Factuality
- Clue-RAG: Towards Accurate and Cost-Efficient Graph-based RAG via Multi-Partite Graph and Query-Driven Iterative Retrieval
- Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember
- Shifting from Ranking to Set Selection for Retrieval Augmented Generation
- The Delta Learning Hypothesis: Preference Tuning on Weak Data can Yield Strong Gains
- HIRAG: Hierarchical-Thought Instruction-Tuning Retrieval-Augmented Generation
- Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution
- Dynamic Injection of Entity Knowledge into Dense Retrievers
- Thinking About Thinking: SAGE-nano's Inverse Reasoning for Self-Aware Language Models
- RAG-R1: Incentivizing the Search and Reasoning Capabilities of LLMs through Multi-query Parallelism
- Learning to Focus: Causal Attention Distillation via Gradient-Guided Token Pruning
- Response Quality Assessment for Retrieval-Augmented Generation via Conditional Conformal Factuality
- EraRAG: Efficient and Incremental Retrieval Augmented Generation for Growing Corpora
- Controlled Retrieval-augmented Context Evaluation for Long-form RAG
- Harnessing the Power of Reinforcement Learning for Language-Model-Based Information Retriever via Query-Document Co-Augmentation
- Deep Research Agents: A Systematic Examination And Roadmap
- Cite Pretrain: Retrieval-Free Knowledge Attribution for Large Language Models
- KAG-Thinker: Interactive Thinking and Deep Reasoning in LLMs via Knowledge-Augmented Generation
- Chain-of-Thought Prompting Obscures Hallucination Cues in Large Language Models: An Empirical Evaluation
- Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs?
- REIS: A High-Performance and Energy-Efficient Retrieval System with In-Storage Processing
- MoR: Better Handling Diverse Queries with a Mixture of Sparse, Dense, and Human Retrievers
- CrEst: Credibility Estimation for Contexts in LLMs via Weak Supervision
- RAGtifier: Evaluating RAG Generation Approaches of State-of-the-Art RAG Systems for the SIGIR LiveRAG Competition
- LTRR: Learning To Rank Retrievers for LLMs
- SimpleDoc: Multi-Modal Document Understanding with Dual-Cue Page Retrieval and Iterative Refinement
- FlexRAG: A Flexible and Comprehensive Framework for Retrieval-Augmented Generation
- Feedback Friction: LLMs Struggle to Fully Incorporate External Feedback
- Constructing and Evaluating Declarative RAG Pipelines in PyTerrier
- Generalization or Hallucination? Understanding Out-of-Context Reasoning in Transformers
- AdaptiveLLM: A Framework for Selecting Optimal Cost-Efficient LLM for Code-Generation Based on CoT Length
- KG-Infused RAG: Augmenting Corpus-Based RAG with External Knowledge Graphs
- Task Matters: Knowledge Requirements Shape LLM Responses to Context-Memory Conflict
- Joint-GCG: Unified Gradient-Based Poisoning Attacks on Retrieval-Augmented Generation Systems
- From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems
- R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning
- ZeroSearch: Incentivize the Search Capability of LLMs without Searching
- LLM-Independent Adaptive RAG: Let the Question Speak for Itself
- Graph-Embedding Empowered Entity Retrieval
- On Generalization across Measurement Systems: LLMs Entail More Test-Time Compute for Underrepresented Cultures
- KARE-RAG: Knowledge-Aware Refinement and Enhancement for RAG
- GraphRAG-Bench: Challenging Domain-Specific Reasoning for Evaluating Graph Retrieval-Augmented Generation
- RewardBench 2: Advancing Reward Model Evaluation
- LEMONADE: A Large Multilingual Expert-Annotated Abstractive Event Dataset for the Real World
- Optimizing the Interface Between Knowledge Graphs and LLMs for Complex Reasoning
- ClueAnchor: Clue-Anchored Knowledge Reasoning Exploration and Optimization for Retrieval-Augmented Generation
- MetaFaith: Faithful Natural Language Uncertainty Expression in LLMs
- Chameleon: A Flexible Data-mixing Framework for Language Model Pretraining and Finetuning
- From Chat Logs to Collective Insights: Aggregative Question Answering
- From Parameters to Prompts: Understanding and Mitigating the Factuality Gap between Fine-Tuned LLMs
- RAGRouter: Learning to Route Queries to Multiple Retrieval-Augmented Language Models
- LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form Generations
- Evaluating the Retrieval Robustness of Large Language Models
- Retrieval-Augmented Generation: A Comprehensive Survey of Architectures, Enhancements, and Robustness Frontiers
- WebDancer: Towards Autonomous Information Seeking Agency
- EvolveSearch: An Iterative Self-Evolving Search Agent
- Faithfulness-Aware Uncertainty Quantification for Fact-Checking the Output of Retrieval Augmented Generation
- How does Misinformation Affect Large Language Model Behaviors and Preferences?
- Pretrained LLMs Learn Multiple Types of Uncertainty
- Prompting is not Enough: Exploring Knowledge Integration and Controllable Generation
- ExAnte: A Benchmark for Ex-Ante Inference in Large Language Models
- KnowTrace: Bootstrapping Iterative Retrieval-Augmented Generation with Structured Knowledge Tracing
- InFact: Informativeness Alignment for Improved LLM Factuality
- CoEvoKG: Co-Evolving Knowledge Graphs with Self-Evolving Search Agents
- Knoll: Creating a Knowledge Ecosystem for Large Language Models
- Removal of Hallucination on Hallucination: Debate-Augmented RAG
- HindSearch: Trajectory-Level Hindsight Critique for Search-Augmented Reinforcement Learning
- GainRAG: Preference Alignment in Retrieval-Augmented Generation through Gain Signal Synthesis
- Why Do Some Inputs Break Low-Bit LLM Quantization?
- Data Mixing Can Induce Phase Transitions in Knowledge Acquisition
- PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning
- Start Classifying: Categorical Critics for LLM Reinforcement Learning
- How Knowledge Popularity Influences and Enhances LLM Knowledge Boundary Perception
- LeTS: Learning to Think-and-Search via Process-and-Outcome Reward Hybridization
- Search Wisely: Mitigating Sub-optimal Agentic Searches By Reducing Uncertainty
- MEGRAG: Multi-Granular Evidence Graphs for Answer-Aware Multi-Hop RAG
- The Rise of Parameter Specialization for Knowledge Storage in Large Language Models
- Walk&Retrieve: Simple Yet Effective Zero-shot Retrieval-Augmented Generation via Knowledge Graph Walks
- O2-Searcher: A Searching-based Agent Model for Open-Domain Open-Ended Question Answering
- FREESON: Retriever-Free Retrieval-Augmented Reasoning via Corpus-Traversing MCTS
- Beyond Hard and Soft: Hybrid Context Compression for Balancing Local and Global Information Retention
- Single LLM, Multiple Roles: A Unified Retrieval-Augmented Generation Framework Using Role-Specific Token Optimization
- Pre-training Limited Memory Language Models with Internal and External Knowledge
- Before Reasoning Can Fail: Pre-Evidence Procedural Failures in Agentic RAG
- Do RAG Systems Really Suffer From Positional Bias?
- LongCat Sparse Attention: Taming the Lightning via Streaming-aware Hierarchical Cross-Layer Indexing
- In-depth Research Impact Summarization through Fine-Grained Temporal Citation Analysis
- Search-GRT: Guided Retrieval Training of Search Agents to Optimize for Complex Question Answering
- s3: You Don't Need That Much Data to Train a Search Agent via RL
- PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent
- The Hallucination Tax of Reinforcement Finetuning
- Process vs. Outcome Reward: Which is Better for Agentic RAG Reinforcement Learning
- LLM-based Query Expansion Fails for Unfamiliar and Ambiguous Queries
- AMAQA: A Metadata-based QA Dataset for RAG Systems
- Accelerating Adaptive Retrieval Augmented Generation via Instruction-Driven Representation Reduction of Retrieval Overlaps
- RAGXplain: From Explainable Evaluation to Actionable Guidance of RAG Pipelines
- BARREL: Boundary-Aware Reasoning for Factual and Reliable LRMs
- Unveiling Knowledge Utilization Mechanisms in LLM-based Retrieval-Augmented Generation
- BiCAA: Bidirectional Credit Assignment for Search-Augmented Agent
- Reinforcing Multi-Turn Reasoning in LLM Agents via Turn-Level Reward Design
- Group-in-Group Policy Optimization for LLM Agent Training
- DRAGON: Domain-specific Robust Automatic Data Generation for RAG Optimization
- Masking in Multi-hop QA: An Analysis of How Language Models Perform with Context Permutation
- Search and Refine During Think: Facilitating Knowledge Refinement for Improved Retrieval-Augmented Reasoning
- Hierarchical Document Refinement for Long-context Retrieval-augmented Generation
- DACL-RAG: Data Augmentation Strategy with Curriculum Learning for Retrieval-Augmented Generation
- Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning
- Scent of Knowledge: Optimizing Search-Enhanced Reasoning with Information Foraging
- Do Large Language Models Know Conflict? Investigating Parametric vs. Non-Parametric Knowledge of LLMs for Conflict Forecasting
- TARGET: Benchmarking Table Retrieval for Generative Tasks
- Answer Presence Drives RAG Rewriting Gains
- Reinforced Internal-External Knowledge Synergistic Reasoning for Efficient Adaptive Search Agent
- Why Uncertainty Estimation Methods Fall Short in RAG: An Axiomatic Analysis
- The Distracting Effect: Understanding Irrelevant Passages in RAG
- References Indeed Matter? Reference-Free Preference Optimization for Conversational Query Reformulation
- Neural Catalog: Scaling Species Recognition with Catalog of Life-Augmented Generation
- SelRoute: Query-Type-Aware Routing for Long-Term Conversational Memory Retrieval
- ContextBudget: Budget-Aware Context Management for Long-Horizon Search Agents
- The Price of Meaning: Why Every Semantic Memory System Forgets
- Agentic Abstention: Do Agents Know When to Stop Instead of Act?
- Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency
- Graph-based Agent Memory: Taxonomy, Techniques, and Applications
- The Model Knows Your Project, Not You: Measuring Recognition in LLMs with NameRank
- Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning
- IslamicLegalBench: Evaluating LLMs Knowledge and Reasoning of Islamic Law Across 1,200 Years of Islamic Pluralist Legal Traditions
- AgentIR: Reasoning-Aware Retrieval for Deep Research Agents
- Truncated Step-Level Sampling with Process Rewards for Retrieval-Augmented Reasoning
- AgentConductor: Topology Evolution for Multi-Agent Competition-Level Code Generation
- Long-Tail Knowledge in Large Language Models: Taxonomy, Mechanisms, Interventions and Implications
- Panini: Continual Learning in Token Space via Structured Memory
- Self-Distilled Agentic Reinforcement Learning
- Co-LMLM: Continuous-Query Limited Memory Language Models
- Can LLMs Introspect? A Reality Check
- NanoKnow: How to Know What Your Language Model Knows
- Calibrate-Then-Act: Cost-Aware Exploration in LLM Agents
- ICL CIPHERS: Quantifying "Learning" in In-Context Learning via Substitution Ciphers
- Conflicts in Texts: Data, Implications and Challenges
- Uncertainty Quantification for Language Models: A Suite of Black-Box, White-Box, LLM Judge, and Ensemble Scorers
- Abstain-R1: Calibrated Abstention and Post-Refusal Clarification via Verifiable RL
- SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
- SE-Search: Self-Evolving Search Agent via Memory and Dense Reward
- Pixel-Grounded Retrieval for Knowledgeable Large Multimodal Models
- Agentic Reinforcement Learning with Self-Distilled Reward Shaping
- Training Documents Reranker with Search Rubrics for Deep Research Agent
- DualRAG: A Dual-Process Approach to Integrate Reasoning and Retrieval for Multi-Hop Question Answering
- PropRAG: Guiding Retrieval with Beam Search over Proposition Paths
- HalluLens: LLM Hallucination Benchmark
- When Compression Scores Cannot Decide: Information Boundaries for Group-Robust LLM Pruning
- A Comprehensive Survey of Knowledge-Based Vision Question Answering Systems: The Lifecycle of Knowledge in Visual Reasoning Task
- Adaptive Orchestration of Modular Generative Information Access Systems
- Credible Plan-Driven RAG Method for Multi-Hop Question Answering
- PURPOSE: Poisoning Conflict Resolution in RAG via Proxy-Fact-Grounded Updates
- IslamicTurathBench: A Multi-Task, Multi-Discipline Benchmark for Evaluating Large Language Models on the Islamic Scholarly Tradition (turath)
- Recall Is Not Enough: A Reader-Context Diagnostic for Budget-Constrained Retrieval-Augmented Generation
- A Survey of Foundation Model-Powered Recommender Systems: From Feature-Based, Generative to Agentic Paradigms
- MIRAGE: A Metric-Intensive Benchmark for Retrieval-Augmented Generation Evaluation
- Multilingual Retrieval-Augmented Generation for Knowledge-Intensive Task
- POLYRAG: Integrating Polyviews into Retrieval-Augmented Generation for Medical Applications
- Retrieval is Not Enough: Enhancing RAG Reasoning through Test-Time Critique and Optimization
- CoLoTa: A Dataset for Entity-based Commonsense Reasoning over Long-Tail Knowledge
- Hypothetical Documents or Knowledge Leakage? Rethinking LLM-based Query Expansion
- Enabling Collaborative Parametric Knowledge Calibration for Retrieval-Augmented Vision Question Answering
- QE-RAG: A Robust Retrieval-Augmented Generation Benchmark for Query Entry Errors
- ACoRN: Noise-Robust Abstractive Compression in Retrieval-Augmented Language Models
- On Linear Representations and Pretraining Data Frequency in Language Models
- AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
- Contextual Information Policy Optimization for Search Agents
- To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling
- DioR: Adaptive Cognitive Detection and Contextual Retrieval Optimization for Dynamic Retrieval-Augmented Generation
- Can We Edit LLMs for Long-Tail Biomedical Knowledge?
- VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents
- Out of Style: RAG's Fragility to Linguistic Variation
- Harnessing the Unseen: The Hidden Influence of Intrinsic Knowledge in Long-Context Language Models
- FamilyTool: A Multi-hop Personalized Tool Use Benchmark
- Knowledge-Instruct: Effective Continual Pre-training from Limited Data using Instructions
Related