A Survey on RAG Meeting LLMs: Towards Retrieval-Augmented Large Language Models
2024/05/10 by Fan, Wenqi, Ding, Yujuan, Ning, Liangbo +5 · 221 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Information Retrieval (cs.IR)
paper · doi:10.48550/arxiv.2405.06211
Abstract
As one of the most advanced techniques in AI, Retrieval-Augmented Generation (RAG) can offer reliable and up-to-date external knowledge, providing huge convenience for numerous tasks. Particularly in the era of AI-Generated Content (AIGC), the powerful capacity of retrieval in providing additional knowledge enables RAG to assist existing generative AI in producing high-quality outputs. Recently, Large Language Models (LLMs) have demonstrated revolutionary abilities in language understanding and generation, while still facing inherent limitations, such as hallucinations and out-of-date internal knowledge. Given the powerful abilities of RAG in providing the latest and helpful auxiliary information, Retrieval-Augmented Large Language Models (RA-LLMs) have emerged to harness external and authoritative knowledge bases, rather than solely relying on the model's internal knowledge, to augment the generation quality of LLMs. In this survey, we comprehensively review existing research studies in RA-LLMs, covering three primary technical perspectives: architectures, training strategies, and applications. As the preliminary knowledge, we briefly introduce the foundations and recent advances of LLMs. Then, to illustrate the practical significance of RAG for LLMs, we systematically review mainstream relevant work by their architectures, training strategies, and application areas, detailing specifically the challenges of each and the corresponding capabilities of RA-LLMs. Finally, to deliver deeper insights, we discuss current limitations and several promising directions for future research. Updated information about this survey can be found at https://advanced-recommender-systems.github.io/RAG-Meets-LLMs/
Cited by
- Problems With Large Language Models for Learner Modelling: Why LLMs Alone Fall Short for Responsible Tutoring in K--12 Education
- MPR-CiteG: Enhancing RAG with Multi-Portfolio Retrieval and Citation-Grounded Generation
- TAMEing Long Contexts in Personalization: Towards Training-Free and State-Aware MLLM Personalized Assistant
- X-GridAgent: An LLM-Powered Agentic AI System for Assisting Power Grid Analysis
- M3KG-RAG: Multi-hop Multimodal Knowledge Graph-enhanced Retrieval-Augmented Generation
- A Large-Language-Model Framework for Automated Humanitarian Situation Reporting
- CIRR: Causal-Invariant Retrieval-Augmented Recommendation with Faithful Explanations under Distribution Shift
- Graph-based Nearest Neighbors with Dynamic Updates via Random Walks
- Scalable Distributed Vector Search via Accuracy Preserving Index Construction
- A systematic assessment of Large Language Models for constructing two-level fractional factorial designs
- Plausibility as Failure: How LLMs and Humans Co-Construct Epistemic Error
- R4: Retrieval-Augmented Reasoning for Vision-Language Models in 4D Spatio-Temporal Space
- Leveraging Spreading Activation for Improved Document Retrieval in Knowledge-Graph-Based RAG Systems
- IaC Generation with LLMs: An Error Taxonomy and A Study on Configuration Knowledge Injection
- AIAuditTrack: A Framework for AI Security system
- Autonomous Construction-Site Safety Inspection Using Mobile Robots: A Multilayer VLM-LLM Pipeline
- From Context to EDUs: Faithful and Structured Context Compression via Elementary Discourse Unit Decomposition
- A Systematic Characterization of LLM Inference on GPUs
- SHRAG: AFrameworkfor Combining Human-Inspired Search with RAG
- HKRAG: Holistic Knowledge Retrieval-Augmented Generation Over Visually-Rich Documents
- HyperbolicRAG: Enhancing Retrieval-Augmented Generation with Hyperbolic Representations
- VecIntrinBench: Benchmarking Cross-Architecture Intrinsic Code Migration for RISC-V Vector
- The Oracle and The Prism: A Decoupled and Efficient Framework for Generative Recommendation Explanation
- WebRec: Enhancing LLM-based Recommendations with Attention-guided RAG from Web
- NeuroPath: Neurobiology-Inspired Path Tracking and Reflection for Semantically Coherent Retrieval
- Grounded by Experience: Generative Healthcare Prediction Augmented with Hierarchical Agentic Retrieval
- BudgetLeak: Membership Inference Attacks on RAG Systems via the Generation Budget Side Channel
- Structured RAG for Answering Aggregative Questions
- Rethinking Retrieval-Augmented Generation for Medicine: A Large-Scale, Systematic Expert Evaluation and Practical Insights
- QuAnTS: Question Answering on Time Series
- Beyond Single Embeddings: Capturing Diverse Targets with Multi-Query Retrieval
- AGRAG: Advanced Graph-based Retrieval-Augmented Generation for LLMs
- A Systematic Literature Review of Code Hallucinations in LLMs: Characterization, Mitigation Methods, Challenges, and Future Directions for Reliable AI
- TreeQA: Enhanced LLM-RAG with logic tree reasoning for reliable and interpretable multi-hop question answering
- Adapting Large Language Models to Emerging Cybersecurity using Retrieval Augmented Generation
- Metacognition Should Be the Scientific Framework for Bounded and Effective Self-Governance in Generative AI
- State of the Art of LLM-Enabled Interaction with Visualization
- ScaleCall -- Agentic Tool Calling at Scale for Fintech: Challenges, Methods, and Deployment Insights
- Metadata-Driven Retrieval-Augmented Generation for Financial Question Answering
- Graph-Guided Concept Selection for Efficient Retrieval-Augmented Generation
- M-Eval: A Heterogeneity-Based Framework for Multi-evidence Validation in Medical RAG Systems
- Dynamically Detect and Fix Hardness for Efficient Approximate Nearest Neighbor Search
- Foundation of Intelligence: Review of Math Word Problems from Human Cognition Perspective
- NeuroGenPoisoning: Neuron-Guided Attacks on Retrieval-Augmented Generation of LLM via Genetic Optimization of External Knowledge
- HA-RAG: Hotness-Aware RAG Acceleration via Mixed Precision and Data Placement
- ResearchGPT: Benchmarking and Training LLMs for End-to-End Computer Science Research Workflows
- FidelityGPT: Correcting Decompilation Distortions with Retrieval Augmented Generation
- Enhancing Hotel Recommendations with AI: LLM-Based Review Summarization and Query-Driven Insights
- AtlasKV: Augmenting LLMs with Billion-Scale Knowledge Graphs in 20GB VRAM
- Comprehending Spatio-temporal Data via Cinematic Storytelling using Large Language Models
- Efficient Toxicity Detection in Gaming Chats: A Comparative Study of Embeddings, Fine-Tuned Transformers and LLMs
- A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Evaluations, and Applications
- Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation
- Refine Thought: A Test-Time Inference Method for Embedding Model Reasoning
- Query-Specific GNN: A Comprehensive Graph Representation Learning Method for Retrieval Augmented Generation
- Autonomous Agents for Scientific Discovery: Orchestrating Scientists, Language, Code, and Physics
- DualResearch: Entropy-Gated Dual-Graph Retrieval for Answer Reconstruction
- RCPU: Rotation-Constrained Error Compensation for Structured Pruning of Large Language Models
- Agentic generative AI for media content discovery at the national football league
- Exposing Citation Vulnerabilities in Generative Engines
- MARS: Co-evolving Dual-System Deep Research via Multi-Agent Reinforcement Learning
- Retrieval-Augmented Code Generation: A Survey with Focus on Repository-Level Approaches
- A Lightweight Large Language Model-Based Multi-Agent System for 2D Frame Structural Analysis
- UNIDOC-BENCH: A Unified Benchmark for Document-Centric Multimodal RAG
- External Data Extraction Attacks against Retrieval-Augmented Large Language Models
- Copy-Paste to Mitigate Large Language Model Hallucinations
- Attribution Gradients: Incrementally Unfolding Citations for Critical Examination of Attributed AI Answers
- MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments
- AutoLabs: Cognitive Multi-Agent Systems with Self-Correction for Autonomous Chemical Experimentation
- Beyond Static Retrieval: Opportunities and Pitfalls of Iterative Retrieval in GraphRAG
- SimulRAG: Simulator-based RAG for Grounding LLMs in Long-form Scientific QA
- SafeSearch: Automated Red-Teaming for the Safety of LLM-Based Search Agents
- LUMINA: Detecting Hallucinations in RAG System with Context-Knowledge Signals
- D-Artemis: A Deliberative Cognitive Framework for Mobile GUI Multi-Agents
- SGMem: Sentence Graph Memory for Long-Term Conversational Agents
- Hierarchical Reranking for Scalable Financial RAG System
- A Knowledge Graph-based Retrieval-Augmented Generation Framework for Algorithm Selection in the Facility Layout Problem
- Revealing Multimodal Causality with Large Language Models
- Less Is More: Elevating RAG via Performance-Driven Context Compression
- Towards the Distributed Large-scale k-NN Graph Construction by Graph Merge
- A Survey on Retrieval And Structuring Augmented Generation with Large Language Models
- Approximate Graph Propagation Revisited: Dynamic Parameterized Queries, Tighter Bounds and Dynamic Updates
- AgentX: Towards Orchestrating Robust Agentic Workflow Patterns with FaaS-hosted MCP Services
- Electricity Demand and Grid Impacts of AI Data Centers: Challenges and Prospects
- Chain or tree? Re-evaluating complex reasoning from the perspective of a matrix of thought
- Lighting the Way for BRIGHT: Reproducible Baselines with Anserini, Pyserini, and RankLLM
- Enhancing Reliability in LLM-Integrated Robotic Systems: A Unified Approach to Security and Safety
- RAG-PRISM: A Personalized, Rapid, and Immersive Skill Mastery Framework with Adaptive Retrieval-Augmented Tutoring
- MSRS: Evaluating Multi-Source Retrieval-Augmented Generation
- Model-Driven Quantum Code Generation Using Large Language Models and Retrieval-Augmented Generation
- LFD: Layer Fused Decoding to Exploit External Knowledge in Retrieval-Augmented Generation
- Diverse And Private Synthetic Datasets Generation for RAG evaluation: A multi-agent framework
- Explicit v.s. Implicit Memory: Exploring Multi-hop Complex Reasoning Over Personalized Information
- Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward
- Retrieval-augmented reasoning with lean language models
- SMA: Who Said That? Auditing Membership Leakage in Semi-Black-box RAG Controlling
- Understanding Users' Privacy Perceptions Towards LLM's RAG-based Memory
- Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs
- Integrating Rules and Semantics for LLM-Based C-to-Rust Translation
- RAGTrace: Understanding and Refining Retrieval-Generation Dynamics in Retrieval-Augmented Generation
- mKG-RAG: Multimodal Knowledge Graph-Enhanced RAG for Visual Question Answering
- Retrieval-Augmented Water Level Forecasting for Everglades
- TURA: Tool-Augmented Unified Retrieval Agent for AI Search
- Method-Based Reasoning for Large Language Models: Extraction, Reuse, and Continuous Improvement
- TreeRanker: Fast and Model-agnostic Ranking System for Code Suggestions in IDEs
- Prompting Large Language Models with Partial Knowledge for Answering Questions with Unseen Entities
- Provably Secure Retrieval-Augmented Generation
- GraphRAG-R1: Graph Retrieval-Augmented Generation with Process-Constrained Reinforcement Learning
- AutoBridge: Automating Smart Device Integration with Centralized Platform
- Fast and Accurate Contextual Knowledge Extraction Using Cascading Language Model Chains and Candidate Answers
- Rote Learning Considered Useful: Generalizing over Memorized Data in LLMs
- Conversations over Clicks: Impact of Chatbots on Information Search in Interdisciplinary Learning
- Graph-Augmented Large Language Model Agents: Current Progress and Future Prospects
- Advancing Shared and Multi-Agent Autonomy in Underwater Missions: Integrating Knowledge Graphs and Retrieval-Augmented Generation
- GREAT: Guiding Query Generation with a Trie for Recommending Related Search about Video at Kuaishou
- CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning
- OmniBench-RAG: A Multi-Domain Evaluation Platform for Retrieval-Augmented Generation Tools
- A Systematic Review of Key Retrieval-Augmented Generation (RAG) Systems: Progress, Gaps, and Future Directions
- Transform Before You Query: A Privacy-Preserving Approach for Vector Retrieval with Embedding Space Alignment
- Stealthy LLM-Driven Data Poisoning Attacks Against Embedding-Based Retrieval-Augmented Recommender Systems
- A Comprehensive Review on Harnessing Large Language Models to Overcome Recommender System Challenges
- DyG-RAG: Dynamic Graph Retrieval-Augmented Generation with Event-Centric Reasoning
- FedRAG: A Framework for Fine-Tuning Retrieval-Augmented Generation Systems
- Evaluating LLMs on Sequential API Call Through Automated Test Generation
- KGRAG-Ex: Explainable Retrieval-Augmented Generation with Knowledge Graph-based Perturbations
- Clue-RAG: Towards Accurate and Cost-Efficient Graph-based RAG via Multi-Partite Graph and Query-Driven Iterative Retrieval
- RAG Safety: Exploring Knowledge Poisoning Attacks to Retrieval-Augmented Generation
- QUEST: Query Optimization in Unstructured Document Analysis
- Rethinking Data Protection in the (Generative) Artificial Intelligence Era
- Scalable evaluation framework for retrieval augmented generation in tobacco research using large Language models
- XGraphRAG: Interactive Visual Analysis for Graph-based Retrieval-Augmented Generation
- Graph-based RAG Enhancement via Global Query Disambiguation and Dependency-Aware Reranking
- MOTOR: Multimodal Optimal Transport via Grounded Retrieval in Medical Visual Question Answering
- FoGE: Fock Space inspired encoding for graph prompting
- Weak-to-Strong GraphRAG: Aligning Weak Retrievers with Large Language Models for Graph-based Retrieval Augmented Generation
- EraRAG: Efficient and Incremental Retrieval Augmented Generation for Growing Corpora
- Engineering RAG Systems for Real-World Applications: Design, Development, and Evaluation
- Knowledge-Aware Diverse Reranking for Cross-Source Question Answering
- Inference Scaled GraphRAG: Improving Multi Hop Question Answering on Knowledge Graphs
- Retrieval-Confused Generation is a Good Defender for Privacy Violation Attack of Large Language Models
- Harnessing the Power of Reinforcement Learning for Language-Model-Based Information Retriever via Query-Document Co-Augmentation
- Deep Research Agents: A Systematic Examination And Roadmap
- Graphs Meet AI Agents: Taxonomy, Progress, and Future Opportunities
- When Does Divide and Conquer Work for Long Context LLM? A Noise Decomposition Framework
- AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding
- KG2QA: Knowledge Graph-enhanced Retrieval-augmented Generation for Communication Standards Question Answering
- SymRAG: Efficient Neuro-Symbolic Retrieval Through Adaptive Query Routing
- Maximally-Informative Retrieval for State Space Model Generation
- StepProof: Step-by-step verification of natural language mathematical proofs
- BioMol-MQA: A Multi-Modal Question Answering Dataset For LLM Reasoning Over Bio-Molecular Interactions
- Joint-GCG: Unified Gradient-Based Poisoning Attacks on Retrieval-Augmented Generation Systems
- Advancing Tool-Augmented Large Language Models via Meta-Verification and Reflection Learning
- MobiEdit: Resource-efficient Knowledge Editing for Personalized On-device LLMs
- V2X-UniPool: Unifying Multimodal Perception and Knowledge Reasoning for Autonomous Driving
- AI Scientists Fail Without Strong Implementation Capability
- Optimizing the Interface Between Knowledge Graphs and LLMs for Complex Reasoning
- CMIE: Combining MLLM Insights with External Evidence for Explainable Out-of-Context Misinformation Detection
- Born a Transformer -- Always a Transformer? On the Effect of Pretraining on Architectural Abilities
- Do LLMs Understand Collaborative Signals? Diagnosis and Repair
- System-driven Cloud Architecture Design Support with Structured State Management and Guided Decision Assistance
- Test-Time Learning for Large Language Models
- CPA-RAG:Covert Poisoning Attacks on Retrieval-Augmented Generation in Large Language Models
- Iterative Self-Incentivization Empowers Large Language Models as Agentic Searchers
- MA-RAG: Multi-Agent Retrieval-Augmented Generation via Collaborative Chain-of-Thought Reasoning
- Direct Retrieval-augmented Optimization: Synergizing Knowledge Selection and Language Models
- Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models
- Real-time Spatial Retrieval Augmented Generation for Urban Environments
- CReSt: A Comprehensive Benchmark for Retrieval-Augmented Generation with Complex Reasoning over Structured Documents
- HENN: A Hierarchical Epsilon Net Navigation Graph for Approximate Nearest Neighbor Search
- Advancing the Scientific Method with Large Language Models: From Hypothesis to Discovery
- Walk&Retrieve: Simple Yet Effective Zero-shot Retrieval-Augmented Generation via Knowledge Graph Walks
- Align-GRAG: Anchor and Rationale Guided Dual Alignment for Graph Retrieval-Augmented Generation
- InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation
- StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
- NExT-Search: Rebuilding User Feedback Ecosystem for Generative AI Search
- Search-GRT: Guided Retrieval Training of Search Agents to Optimize for Complex Question Answering
- SCAN: Semantic Document Layout Analysis for Textual and Visual Retrieval-Augmented Generation
- Process vs. Outcome Reward: Which is Better for Agentic RAG Reinforcement Learning
- Divide by Question, Conquer by Agent: SPLIT-RAG with Question-Driven Graph Partitioning
- Know3-RAG: A Knowledge-aware RAG Framework with Adaptive Retrieval, Generation, and Filtering
- Real-Time Hybrid Retrieval in Hyperbolic Space for Retrieval-Augmented Generation on Edge Devices
- Demystifying and Enhancing the Efficiency of Large Language Model Based Search Agents
- Supporting Cybersecurity Risk Management for Medical Devices via the SECUMAN Ontology and Shapes
- SubGCache: Accelerating Graph-based RAG with Subgraph-level KV Cache
- mmRAG: A Modular Benchmark for Retrieval-Augmented Generation over Text, Tables, and Knowledge Graphs
- Chatting with Papers: A Hybrid Approach Using LLMs and Knowledge Graphs
- Towards Automated Situation Awareness: A RAG-Based Framework for Peacebuilding Reports
- Benchmarking Retrieval-Augmented Generation for Chemistry
- Towards AI-Driven Human-Machine Co-Teaming for Adaptive and Agile Cyber Security Operation Centers
- KG-HTC: Integrating Knowledge Graphs into LLMs for Effective Zero-shot Hierarchical Text Classification
- Retrieval-Augmented Generation in Biomedicine: A Survey of Technologies, Datasets, and Clinical Applications
- Swarm: Co-Activation Aware KVCache Offloading Across Multiple SSDs
- Distributed Retrieval-Augmented Generation
- RAG over Thinking Traces Can Improve Reasoning Tasks
- MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image
- How Do LLMs Encode Scientific Quality? An Empirical Study Using Monosemantic Features from Sparse Autoencoders
- Feedback by Design: Understanding and Overcoming User Feedback Barriers in Conversational Agents
- Making Theft Useless: Adulteration-Based Protection of Proprietary Knowledge Graphs in GraphRAG Systems
- Towards Large Language Models for Lunar Mission Planning and In Situ Resource Utilization
- Reconstructing Context: Evaluating Advanced Chunking Strategies for Retrieval-Augmented Generation
- PriHA: A RAG-Enhanced LLM Framework for Primary Healthcare Assistant in Hong Kong
- Bridging Instead of Replacing Online Coding Communities with AI through Community-Enriched Chatbot Designs
- TumorBoard: Evidence-Grounded Multi-Agent Decision Support for Longitudinal Neuro-Oncology
- LLM-assisted Graph-RAG Information Extraction from IFC Data
- Retrieval Augmented Generation Evaluation in the Era of Large Language Models: A Comprehensive Survey
- Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval
- Beyond Misinformation: A Conceptual Framework for Studying AI Hallucinations in (Science) Communication
- Everything You Wanted to Know About LLM-based Vulnerability Detection But Were Afraid to Ask
- Diffusion Generative Recommendation with Continuous Tokens
- Exploring Backdoor Attack and Defense for LLM-empowered Recommendations
- ReZero: Enhancing LLM search ability by trying one-more-time
- A Survey of Personalization: From RAG to Agent
- ControlNET: A Firewall for RAG-based LLM System
- VectorLiteRAG: Latency-Aware and Fine-Grained Resource Partitioning for Efficient RAG
- RAG-VR: Leveraging Retrieval-Augmented Generation for 3D Question Answering in VR Environments
- The Multi-Round Diagnostic RAG Framework for Emulating Clinical Reasoning
- Toward Holistic Evaluation of Recommender Systems Powered by Generative Models
- Simplifying Data Integration: SLM-Driven Systems for Unified Semantic Queries Across Heterogeneous Databases
- Retrieval Augmented Generation with Collaborative Filtering for Personalized Text Generation
- SciSciGPT: Advancing Human-AI Collaboration in the Science of Science
- RS-RAG: Bridging Remote Sensing Imagery and Comprehensive Knowledge with a Multi-Modal Dataset and Retrieval-Augmented Generation Model
Related