Benchmarking Large Language Models in Retrieval-Augmented Generation
2023/09/04 by Jiawei Chen, Hongyu Lin, Chen, Jiawei +5 · 1 voice · 103 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling #cs.CL
paper · pdf · doi:10.48550/arxiv.2309.01431
openalex publication_date 2023/09/04 · arxiv published 2023/09/04 · arxiv updated 2023/12/20 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Retrieval-Augmented Generation (RAG) is a promising approach for mitigating the hallucination of large language models (LLMs). However, existing research lacks rigorous evaluation of the impact of retrieval-augmented generation on different large language models, which make it challenging to identify the potential bottlenecks in the capabilities of RAG for different LLMs. In this paper, we systematically investigate the impact of Retrieval-Augmented Generation on large language models. We analyze the performance of different large language models in 4 fundamental abilities required for RAG, including noise robustness, negative rejection, information integration, and counterfactual robustness. To this end, we establish Retrieval-Augmented Generation Benchmark (RGB), a new corpus for RAG evaluation in both English and Chinese. RGB divides the instances within the benchmark into 4 separate testbeds based on the aforementioned fundamental abilities required to resolve the case. Then we evaluate 6 representative LLMs on RGB to diagnose the challenges of current LLMs when applying RAG. Evaluation reveals that while LLMs exhibit a certain degree of noise robustness, they still struggle significantly in terms of negative rejection, information integration, and dealing with false information. The aforementioned assessment outcomes indicate that there is still a considerable journey ahead to effectively apply RAG to LLMs.
Cited by
- Detecting Hallucinations in Graph Retrieval-Augmented Generation via Attention Patterns and Semantic Alignment
- A Systematic Framework for Enterprise Knowledge Retrieval: Leveraging LLM-Generated Metadata to Enhance RAG Systems
- Rethinking Retrieval: From Traditional Retrieval Augmented Generation to Agentic and Non-Vector Reasoning Systems in the Financial Domain for Large Language Models
- CARE-RAG - Clinical Assessment and Reasoning in RAG
- Noise-Robust Abstractive Compression in Retrieval-Augmented Language Models
- GRAPH-GRPO-LEX: Contract Graph Modeling and Reinforcement Learning with Group Relative Policy Optimization
- RAGalyst: Automated Human-Aligned Agentic Evaluation for Domain-Specific RAG
- Plan of Knowledge: Retrieval-Augmented Large Language Models for Temporal Knowledge Graph Question Answering
- Retrieval Augmented Generation (RAG) for Fintech: Agentic Design and Evaluation
- Metadata-Driven Retrieval-Augmented Generation for Financial Question Answering
- A Video Is Not Worth a Thousand Words
- A Feasibility Study on Usability and Trust among Population Groups of a Medical Avatar Supported by Large Language Models with Retrieval Augmented Generation
- WebSeer: Training Deeper Search Agents through Reinforcement Learning with Self-Reflection
- AcademicEval: Live Long-Context LLM Benchmark
- Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations
- Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool Usage
- When Retrieval Succeeds and Fails: Rethinking Retrieval-Augmented Generation for LLMs
- Large Language Models: A Paradigm Shift for Dementia Diagnosis and Care
- Towards an Efficient, Customizable, and Accessible AI Tutor
- GRAD: Generative Retrieval-Aligned Demonstration Sampler for Efficient Few-Shot Reasoning
- Incentive-Aligned Multi-Source LLM Summaries
- Enhancing LLM-based Fault Localization with a Functionality-Aware Retrieval-Augmented Generation Framework
- OptGraph: Large Language Models Enhanced Evolutionary Optimization Via Graph Retrieval-Augmented Generation
- ContextNest: Verifiable Context Governance for Autonomous AI Agent
- RELATE: Relation Extraction in Biomedical Abstracts with LLMs and Ontology Constraints
- WildClaims: Information Access Conversations in the Wild(Chat)
- Cuckoo Attack: Stealthy and Persistent Attacks Against AI-IDE
- CARGO: A Framework for Confidence-Aware Routing of Large Language Models
- Linguistic Nepotism: Trading-off Quality for Language Preference in Multilingual RAG
- Who Taught the Lie? Responsibility Attribution for Poisoned Knowledge in Retrieval-Augmented Generation
- A Dynamic Knowledge Update-Driven Model with Large Language Models for Fake News Detection
- HiChunk: Evaluating and Enhancing Retrieval-Augmented Generation with Hierarchical Chunking
- RAG-PRISM: A Personalized, Rapid, and Immersive Skill Mastery Framework with Adaptive Retrieval-Augmented Tutoring
- LFD: Layer Fused Decoding to Exploit External Knowledge in Retrieval-Augmented Generation
- Diverse And Private Synthetic Datasets Generation for RAG evaluation: A multi-agent framework
- Can we Evaluate RAGs with Synthetic Data?
- Leveraging LLMs for Smart Cities Qualitative Data Analysis
- RAGTrace: Understanding and Refining Retrieval-Generation Dynamics in Retrieval-Augmented Generation
- Are We on the Right Way for Assessing Document Retrieval-Augmented Generation?
- ActionSink: Toward Precise Robot Manipulation with Dynamic Integration of Action Flow
- CTTS: Collective Test-Time Scaling
- The SMeL Test: A simple benchmark for media literacy in language models
- CoGrader: Transforming Instructors' Assessment of Project Reports through Collaborative LLM Integration
- QSAF: A Novel Mitigation Framework for Cognitive Degradation in Agentic AI
- A Systematic Review of Key Retrieval-Augmented Generation (RAG) Systems: Progress, Gaps, and Future Directions
- Small Data Explainer -- The impact of small data methods in everyday life
- Weak-to-Strong GraphRAG: Aligning Weak Retrievers with Large Language Models for Graph-based Retrieval Augmented Generation
- The Next Phase of Scientific Fact-Checking: Advanced Evidence Retrieval from Complex Structured Academic Papers
- Keeping Medical AI Healthy and Trustworthy: A Review of Detection and Correction Methods for System Degradation
- REIS: A High-Performance and Energy-Efficient Retrieval System with In-Storage Processing
- Large Language Model Empowered Design of Fluid Antenna Systems: Challenges, Frameworks, and Case Studies for 6G
- Evaluating and Improving Robustness in Large Language Models: A Survey and Future Directions
- Watermarking LLM-Generated Datasets in Downstream Tasks
- MALM: A Multi-Information Adapter for Large Language Models to Mitigate Hallucination
- Defending against Indirect Prompt Injection by Instruction Detection
- Locality Preserving Markovian Transition for Instance Retrieval
- CDE-Mapper: Using Retrieval-Augmented Language Models for Linking Clinical Data Elements to Controlled Vocabularies
- Magic Mushroom: A Customizable Benchmark for Fine-grained Analysis of Retrieval Noise Erosion in RAG Systems
- RARE: Retrieval-Aware Robustness Evaluation for Retrieval-Augmented Generation Systems
- LogiDebrief: A Signal-Temporal Logic based Automated Debriefing Approach with Large Language Models Integration
- RAGRouter: Learning to Route Queries to Multiple Retrieval-Augmented Language Models
- VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning
- Do LLMs Understand Collaborative Signals? Diagnosis and Repair
- A Lightweight Multi-Expert Generative Language Model System for Engineering Information and Knowledge Extraction
- CPA-RAG:Covert Poisoning Attacks on Retrieval-Augmented Generation in Large Language Models
- Knoll: Creating a Knowledge Ecosystem for Large Language Models
- AI-Driven Climate Policy Scenario Generation for Sub-Saharan Africa
- Benchmarking Poisoning Attacks against Retrieval-Augmented Generation
- CReSt: A Comprehensive Benchmark for Retrieval-Augmented Generation with Complex Reasoning over Structured Documents
- Walk&Retrieve: Simple Yet Effective Zero-shot Retrieval-Augmented Generation via Knowledge Graph Walks
- InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation
- SciCUEval: A Comprehensive Dataset for Evaluating Scientific Context Understanding in Large Language Models
- Transparent and Robust RAG: Adaptive-Reward Reinforcement Learning for Decision Traceability
- CAIM: Development and Evaluation of a Cognitive AI Memory Framework for Long-Term Interaction with Intelligent Agents
- mmRAG: A Modular Benchmark for Retrieval-Augmented Generation over Text, Tables, and Knowledge Graphs
- RAG-TESTER: Automated End-to-End Testing of Retrieval-Augmented Large Language Models
- Optimizing Retrieval-Augmented Generation: Analysis of Hyperparameter Impact on Performance and Efficiency
- Towards Requirements Engineering for RAG Systems
- DynamicRAG: Leveraging Outputs of Large Language Model as Feedback for Dynamic Reranking in Retrieval-Augmented Generation
- Distributed Retrieval-Augmented Generation
- ChronoGrapher: Event-Centric Knowledge Graph Construction via Informed Graph Traversal
- IslamicLegalBench: Evaluating LLMs Knowledge and Reasoning of Islamic Law Across 1,200 Years of Islamic Pluralist Legal Traditions
- Benchmark Test-Time Scaling of General LLM Agents
- Traceback of Poisoning Attacks to Retrieval-Augmented Generation
- UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities
- Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets
- Enhancing Speech-to-Speech Dialogue Modeling with End-to-End Retrieval-Augmented Generation
- Designing Digital Humans with Ambient Intelligence
- Assessing the Potential of Generative Agents in Crowdsourced Fact-Checking
- Practical Poisoning Attacks against Retrieval-Augmented Generation
- Compass-V2 Technical Report
- The Great Nugget Recall: Automating Fact Extraction and RAG Evaluation with Large Language Models
- Retrieval Augmented Generation Evaluation in the Era of Large Language Models: A Comprehensive Survey
- PROMPTEVALS: A Dataset of Assertions and Guardrails for Custom Production Large Language Model Pipelines
- LegalRAG: A Hybrid RAG System for Multilingual Legal Information Retrieval
- QE-RAG: A Robust Retrieval-Augmented Generation Benchmark for Query Entry Errors
- ACoRN: Noise-Robust Abstractive Compression in Retrieval-Augmented Language Models
- Retrieval-Augmented Generation with Conflicting Evidence
- Shared Disk KV Cache Management for Efficient Multi-Instance Inference in RAG-Powered LLMs
- Cleo: A Transparent and Controllable Chatbot for Conversational Commerce
- CRAB: A Benchmark for Evaluating Curation of Retrieval-Augmented LLMs in Biomedicine
- Out of Style: RAG's Fragility to Linguistic Variation
- PR-Attack: Coordinated Prompt-RAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel Optimization
Discussions
Related