Towards Agentic RAG with Deep Reasoning: A Survey of RAG-Reasoning Systems in LLMs
2025/07/13 by Li, Yangning, Zhang, Weizhi, Yang, Yuyao +17 · 16 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2507.09477
Abstract
Retrieval-Augmented Generation (RAG) lifts the factuality of Large Language Models (LLMs) by injecting external knowledge, yet it falls short on problems that demand multi-step inference; conversely, purely reasoning-oriented approaches often hallucinate or mis-ground facts. This survey synthesizes both strands under a unified reasoning-retrieval perspective. We first map how advanced reasoning optimizes each stage of RAG (Reasoning-Enhanced RAG). Then, we show how retrieved knowledge of different type supply missing premises and expand context for complex inference (RAG-Enhanced Reasoning). Finally, we spotlight emerging Synergized RAG-Reasoning frameworks, where (agentic) LLMs iteratively interleave search and reasoning to achieve state-of-the-art performance across knowledge-intensive benchmarks. We categorize methods, datasets, and open challenges, and outline research avenues toward deeper RAG-Reasoning systems that are more effective, multimodally-adaptive, trustworthy, and human-centric. The collection is available at https://github.com/DavidZWZ/Awesome-RAG-Reasoning.
Citations
- From Web Search towards Agentic Deep Research: Incentivizing Search with Reasoning Agents
- ZeroSearch: Incentivize the Search Capability of LLMs without Searching
- LLM-Independent Adaptive RAG: Let the Question Speak for Itself
- SymbioticRAG: Enhancing Document Intelligence Through Human-LLM Symbiotic Collaboration
- WebThinker: Empowering Large Reasoning Models with Deep Research Capability
- BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese
- Generative to Agentic AI: Survey, Conceptualization, and Challenges
- DualRAG: A Dual-Process Approach to Integrate Reasoning and Retrieval for Multi-Hop Question Answering
- Credible Plan-Driven RAG Method for Multi-Hop Question Answering
- BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
- ARise: Towards Knowledge-Augmented Reasoning via Risk-Adaptive Search
- HM-RAG: Hierarchical Multi-Agent Multimodal Retrieval Augmented Generation
- Collab-RAG: Boosting Retrieval-Augmented Generation for Complex Question Answering via White-Box and Black-Box LLM Collaboration
- ReaRAG: Knowledge-guided Reasoning Enhances Factuality of Large Reasoning Models with Iterative Retrieval Augmented Generation
- MCTS-RAG: Enhancing Retrieval-Augmented Generation with Monte Carlo Tree Search
- Open Deep Search: Democratizing Search with Open-source Reasoning Agents
- Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
- SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
- MDocAgent: A Multi-Modal Multi-Agent Framework for Document Understanding
- RAG-RL: Advancing Retrieval-Augmented Generation via RL and Curriculum Learning
- Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
- RAG-KG-IL: A Multi-Agent Hybrid Framework for Reducing Hallucinations and Enhancing LLM Reasoning through RAG and Incremental Knowledge Graph Learning Integration
- SurgRAW: Multi-Agent Workflow with Chain-of-Thought Reasoning for Surgical Intelligence
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- OpenRAG: Optimizing RAG End-to-End via In-Context Retrieval Learning
- LightThinker: Thinking Step-by-Step Compression
- TokenSkip: Controllable Chain-of-Thought Compression in LLMs
- Self-Training Large Language Models for Tool-Use Without Demonstrations
- ReARTeR: Retrieval-Augmented Reasoning with Trustworthy Process Rewarding
- WebWalker: Benchmarking LLMs in Web Traversal
- Search-o1: Agentic Search-Enhanced Large Reasoning Models
- Agent Laboratory: Using LLM Agents as Research Assistants
- Tree-based RAG-Agent Recommendation System: A Case Study in Medical Test Data
- Cold-Start Recommendation towards the Era of Large Language Models (LLMs): A Comprehensive Survey and Roadmap
- LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating
- MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge
- PERC: Plan-As-Query Example Retrieval for Underrepresented Code Generation
- RAG-Star: Enhancing Deliberative Reasoning with Retrieval Augmented Verification and Refinement
- A Collaborative Multi-Agent Approach to Retrieval-Augmented Generation Across Diverse Data
- SLA Management in Reconfigurable Multi-Agent RAG: A Systems Approach to Question Answering
- Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent
- Eliciting Critical Reasoning in Retrieval-Augmented Language Models via Contrastive Explanations
- CoPS: Empowering LLM Agents with Provable Cross-Task Experience Sharing
- RuleRAG: Rule-Guided Retrieval-Augmented Generation with Language Models for Question Answering
- StructRAG: Boosting Knowledge Intensive Reasoning of LLMs via Inference-time Hybrid Information Structurization
- GIVE: Structured Reasoning of Large Language Models with Knowledge Graph Inspired Veracity Extrapolation
- Retrieval-Augmented Decision Transformer: External Memory for In-context RL
- ALR2: A Retrieve-then-Reason Framework for Long-context Question Answering
- Retrieval Augmented Generation (RAG) and Beyond: A Comprehensive Survey on How to Make your LLMs use External Data More Wisely
- Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation
- SFR-RAG: Towards Contextually Faithful LLMs
- Retrieval-Augmented Hierarchical in-Context Reinforcement Learning and Hindsight Modular Reflections for Task Planning with LLMs
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- Re-Invoke: Tool Invocation Rewriting for Zero-Shot Tool Retrieval
- Improving Retrieval Augmented Language Model with Self-Reasoning
- MindSearch: Mimicking Human Minds Elicits Deep AI Searcher
- MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
- BeamAggR: Beam Aggregation Reasoning over Multi-source Knowledge for Multi-hop Question Answering
- UDA: A Benchmark Suite for Retrieval Augmented Generation in Real-world Document Analysis
- Learning to Plan for Retrieval-Augmented Large Language Models from Knowledge Graphs
- Retrieval Meets Reasoning: Dynamic In-Context Editing for Long-Text Understanding
- AvaTaR: Optimizing LLM Agents for Tool Usage via Contrastive Reasoning
- VeraCT Scan: Retrieval-Augmented Fake News Detection with Justifiable Reasoning
- RATT: A Thought Structure for Coherent and Correct LLM Reasoning
- Chain of Agents: Large Language Models Collaborating on Long-Context Tasks
- MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
- GNN-RAG: Graph Neural Retrieval for Large Language Model Reasoning
- M-RAG: Reinforcing Large Language Model Performance through Retrieval-Augmented Generation with Multiple Partitions
- DALK: Dynamic Co-Augmentation of LLMs and KG to answer Alzheimer's Disease Questions with Scientific Literature
- Recall, Retrieve and Reason: Towards Better In-Context Relation Extraction
- CBR-RAG: Case-Based Reasoning for Retrieval Augmented Generation in LLMs for Legal Question Answering
- Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity
- RAFT: Adapting Language Model to Domain Specific RAG
- Re-Search for The Truth: Multi-round Retrieval-augmented Large Language Models are Strong Fake News Detectors
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- RAT: Retrieval Augmented Thoughts Elicit Context-Aware Reasoning in Long-Horizon Generation
- MATHSENSEI: A Tool-Augmented Large Language Model for Mathematical Reasoning
- SciAgent: Tool-augmented Language Models for Scientific Reasoning
- RAP: Retrieval-Augmented Planning with Contextual Memory for Multimodal LLM Agents
- MultiHop-RAG: Benchmarking Retrieval-Augmented Generation for Multi-Hop Queries
- INTERS: Unlocking the Power of Large Language Models in Search with Instruction Tuning
- IAG: Induction-Augmented Generation Framework for Answering Reasoning Questions
- GAIA: a benchmark for General AI Assistants
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Chain-of-Note: Enhancing Robustness in Retrieval-Augmented Language Models
- JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models
- A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
- InstructRetro: Instruction Tuning post Retrieval-Augmented Pretraining
- GROVE: A Retrieval-augmented Complex Story Generation Framework with A Forest of Evidence
- Making Retrieval-Augmented Language Models Robust to Irrelevant Context
- RA-DIT: Retrieval-Augmented Dual Instruction Tuning
- Knowledge Graph Prompting for Multi-Document Question Answering
- Counterfactually Auditable Lifecycle Certification for Autonomous Agents
- Think-on-Graph: Deep and Responsible Reasoning of Large Language Model on Knowledge Graph
- Dr.ICL: Demonstration-Retrieved In-context Learning
- Making Language Models Better Tool Learners with Execution Feedback
- ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings
- UPRISE: Universal Prompt Retrieval for Improving Zero-Shot Evaluation
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions
- Measuring and Narrowing the Compositionality Gap in Language Models
- Distributing Accountability, Not Capability: Phase Separation and the LLM Workflow Quadrant in Autonomous AI Agent Architectures
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
- Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
- WebGPT: Browser-assisted question-answering with human feedback
- QuALITY: Question Answering with Long Input Texts, Yes!
- TopiOCQA: Open-domain Conversational Question Answering with Topic Switching
- CREAK: A Dataset for Commonsense Reasoning over Entity Knowledge
- MiniF2F: a cross-system benchmark for formal Olympiad-level mathematics
- QA-GNN: Reasoning with Language Models and Knowledge Graphs for Question Answering
- Measuring Mathematical Problem Solving With the MATH Dataset
- Did Aristotle Use a Laptop? A Question Answering Benchmark with Implicit Reasoning Strategies
- Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps
- IIRC: A Dataset of Incomplete Information Reading Comprehension Questions
- Explainable Automated Fact-Checking for Public Health Claims
- ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
- RefactoringMiner 2.0
- BIGPATENT: A Large-Scale Dataset for Abstractive and Coherent Summarization
- PullNet: Open Domain Question Answering with Iterative Retrieval on Knowledge Bases and Text
- HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
- Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional\n Neural Networks for Extreme Summarization
- CrisisMMD: Multimodal Twitter Datasets from Natural Disasters
- The Web as a Knowledge-base for Answering Complex Questions
- FEVER: a large-scale dataset for Fact Extraction and VERification
- DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset
- Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems
- TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for\n Reading Comprehension
- Retrieval is Not Enough: Enhancing RAG Reasoning through Test-Time Critique and Optimization
- Talk to Right Specialists: Iterative Routing in Multi-agent Systems for Question Answering
- PersonaAgent: Bridging Memory and Action for Personalized LLM Agents
- Supervising the search process produces reliable and generalizable information-seeking agents
Cited by
Related