EviNote-RAG: Enhancing RAG Models via Answer-Supportive Evidence Notes
2025/08/31 by Dai, Yuqin, Guoqing Wang, Yuan Wang +28 · 5 citations
Computer Science · #Computation and Language (cs.CL) #Expert finding and Q&A systems #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2509.00877
openalex publication_date 2025/08/31 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Retrieval-Augmented Generation (RAG) has advanced open-domain question answering by incorporating external information into model reasoning. However, effectively leveraging external information to enhance reasoning presents the following challenges: (1) low signal-to-noise ratio, where answer-supportive external information is diluted by irrelevant material, and (2) error accumulation, which arises in multi-hop reasoning when incomplete or misleading information is incorporated. To address these challenges, we introduce EviNote-RAG, a framework that follows a retrieve-note-answer workflow. Instead of reasoning directly over raw external information, the model first produces Supportive-Evidence Notes (SENs), which concisely preserve answer-critical information and explicitly mark key and uncertainty information to improve accuracy. We further design an entailment-based Evidence Quality Reward (EQR) to ensure that SENs are logically sufficient to derive the final answer, thereby enhancing SENs' quality. Experiments on both in-domain and out-of-domain QA benchmarks show that EviNote-RAG achieves state-of-the-art performance, improving answer accuracy, training stability, robustness, and efficiency. In particular, it yields relative F1 gains of 20% on HotpotQA (+0.093), 40% on Bamboogle (+0.151), and 91% on 2Wiki (+0.256), benefiting from improvements in the reasoning process.
Citations
- Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward
- Careful Queries, Credible Results: Teaching RAG Models Advanced Web Search Tools with Reinforcement Learning
- R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning
- R1-Searcher++: Incentivizing the Dynamic Knowledge Acquisition of LLMs via Reinforcement Learning
- WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning
- ZeroSearch: Incentivize the Search Capability of LLMs without Searching
- ReasonIR: Training Retrievers for Reasoning Tasks
- Synergizing RAG and Reasoning: A Systematic Review
- DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments
- ReaRAG: Knowledge-guided Reasoning Enhances Factuality of Large Reasoning Models with Iterative Retrieval Augmented Generation
- Open Deep Search: Democratizing Search with Open-source Reasoning Agents
- RAG-RL: Advancing Retrieval-Augmented Generation via RL and Curriculum Learning
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Beyond Outlining: Heterogeneous Recursive Planning for Adaptive Long-form Writing with Language Models
- R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
- From RAG to Memory: Non-Parametric Continual Learning for Large Language Models
- Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools
- MedRAG: Enhancing Retrieval-augmented Generation with Knowledge Graph-Elicited Reasoning for Healthcare Copilot
- DeepRAG: Thinking to Retrieve Step by Step for Large Language Models
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- PIKE-RAG: sPecIalized KnowledgE and Rationale Augmented Generation
- ReARTeR: Retrieval-Augmented Reasoning with Trustworthy Process Rewarding
- Search-o1: Agentic Search-Enhanced Large Reasoning Models
- Qwen2.5 Technical Report
- An Agent Framework for Real-Time Financial Information Searching with Large Language Models
- RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models
- WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning
- Plan*RAG: Efficient Test-Time Planning for Retrieval Augmented Generation
- Long-Context LLMs Meet RAG: Overcoming Challenges for Long Inputs in RAG
- Inference Scaling for Long-Context Retrieval Augmented Generation
- Can We Further Elicit Reasoning in LLMs? Critic-Guided Planning with Retrieval-Augmentation for Solving Challenging Tasks
- A Compressive Memory-based Retrieval Approach for Event Argument Extraction
- Into the Unknown Unknowns: Engaged Human Learning through Participation in Language Model Agent Conversations
- Hierarchical Retrieval-Augmented Generation Model with Rethink for Multi-hop Question Answering
- Summary of a Haystack: A Challenge to Long-Context LLMs and RAG Systems
- PlanRAG: A Plan-then-Retrieval Augmented Generation for Generative Large Language Models as Decision Makers
- RE-RAG: Improving Open-Domain QA Performance and Interpretability with Relevance Estimator in Retrieval-Augmented Generation
- Chain of Agents: Large Language Models Collaborating on Long-Context Tasks
- Don't Forget to Connect! Improving RAG with Graph-based Reranking
- SuRe: Summarizing Retrievals using Answer Candidates for Open-domain QA of LLMs
- Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity
- RAFT: Adapting Language Model to Domain Specific RAG
- Metacognitive Retrieval-Augmented Large Language Models
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
- LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression
- RECOMP: Improving Retrieval-Augmented LMs with Compression and Selective Augmentation
- Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
- 🧜Siren’s Song in the AI Ocean: A Survey on Hallucination in Large Language Models
- Ghostbuster: Detecting Text Ghostwritten by Large Language Models
- Enhancing Retrieval-Augmented Large Language Models with Iterative Retrieval-Generation Synergy
- Active Retrieval Augmented Generation
- Large Language Models Can Be Easily Distracted by Irrelevant Context
- Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions
- Text Embeddings by Weakly-Supervised Contrastive Pre-training
- Query Refinement Prompts for Closed-Book Long-Form Question Answering
- Measuring and Narrowing the Compositionality Gap in Language Models
- Distributing Accountability, Not Capability: Phase Separation and the LLM Workflow Quadrant in Autonomous AI Agent Architectures
- Re2G: Retrieve, Rerank, Generate
- PaLM: Scaling Language Modeling with Pathways
- Survey of Hallucination in Natural Language Generation
- WebGPT: Browser-assisted question-answering with human feedback
- Evidentiality-guided Generation for Knowledge-Intensive NLP Tasks
- MuSiQue: Multihop Questions via Single-hop Question Composition
- Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps
- Language Models are Few-Shot Learners
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Dense Passage Retrieval for Open-Domain Question Answering
- REALM: Retrieval-Augmented Language Model Pre-Training
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
- HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
- TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for\n Reading Comprehension
- Retrieval is Not Enough: Enhancing RAG Reasoning through Test-Time Critique and Optimization
- Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
- When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories
Cited by
Related