Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
2025/03/12 by Bowen Jin, Jin, Bowen, Hansi Zeng +14 · 14 voices · 468 citations
Computer Science · #Information Retrieval and Search Behavior #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.IR
paper · pdf · doi:10.48550/arxiv.2503.09516
openalex publication_date 2025/03/12 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Efficiently acquiring external knowledge and up-to-date information is essential for effective reasoning and text generation in large language models (LLMs). Prompting advanced LLMs with reasoning capabilities to use search engines during inference is often suboptimal, as the LLM might not fully possess the capability on how to interact optimally with the search engine. This paper introduces Search-R1, an extension of reinforcement learning (RL) for reasoning frameworks where the LLM learns to autonomously generate (multiple) search queries during step-by-step reasoning with real-time retrieval. Search-R1 optimizes LLM reasoning trajectories with multi-turn search interactions, leveraging retrieved token masking for stable RL training and a simple outcome-based reward function. Experiments on seven question-answering datasets show that Search-R1 improves performance by 41% (Qwen2.5-7B) and 20% (Qwen2.5-3B) over various RAG baselines under the same setting. This paper further provides empirical insights into RL optimization methods, LLM choices, and response length dynamics in retrieval-augmented reasoning. The code and model checkpoints are available at https://github.com/PeterGriffinJin/Search-R1.
Citations
Cited by
- ARCO: Adaptive Rubrics with Co-Evolution for Multi-Step LLM-Based Agents
- From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search
- InteractComp: Evaluating Search Agents With Ambiguous Queries
- Learning to Reason for Factuality
- PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning
- ToolAnchor: Anchoring Counterfactual Context to Boost Agentic Tool-use Capability
- In-the-Flow Agentic System Optimization for Effective Planning and Tool Use
- Why Does Feedback-Augmented Self-Distillation Fail to Improve Retrieval-Interleaved Search Agents?
- Efficient Multi-round LLM Inference over Disaggregated Serving
- Dr. Zero: Self-Evolving Search Agents without Training Data
- Search-on-Graph-R1: Training Large Language Models to Search Knowledge Graphs with Reinforcement Learning
- DRNOISE: Benchmarking Deep Research Agents in Misleading Evidence Environments
- UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation
- Multi-Turn On-Policy Distillation with Prefix Replay
- SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
- LOTAPO: Leave-One-Turn Attribution for Self-Generated Process Rewards in Multi-Turn Search Reasoning
- SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration
- From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training
- AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge Graphs
- EvoSQL: Memory-Augmented Critic-Generator Co-Evolution for Text-to-SQL
- Superintelligent Retrieval Agent: The Next Frontier of Agentic Retrieval
- Eta Given Delta: Defining LLM Tool Efficiency With Marginal Tool Utility
- AI Can Learn Scientific Taste
- Meta-Reinforcement Learning with Self-Reflection for Agentic Search
- The Reasoning Trap: How Enhancing LLM Reasoning Amplifies Tool Hallucination
- R-Zero: Self-Evolving Reasoning LLM from Zero Data
- DeepRAG: Thinking to Retrieve Step by Step for Large Language Models
- TimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinforcement Learning
- Agentic Entropy-Balanced Policy Optimization
- Dynamic Vocabulary Pruning: Stable LLM-RL by Taming the Tail
- Video-BrowseComp: Benchmarking Agentic Video Research on Open Web
- FoldAct: Efficient and Stable Context Folding for Long-Horizon Search Agents
- RollArt: Disaggregated Multi-Task Agentic RL Training at Scale
- Role-Based Fault Tolerance System for LLM RL Post-Training
- Towards a Relevance Posterior in Neural Information Access
- A New Role for Relevance: Guiding Corpus Interaction in Agentic Search
- Co-Evolving Graph and Text Memory for Training-Free Multi-Hop Question Answering
- EviBack: Search-Agent Reinforcement Learning via Evidence-Constrained Teacher Backoff
- SearchArt: Training Long-Horizon Search Agent with Scalable Synthetic and Verified Task
- ATOD: Annealed Turn-Aware On-Policy Distillation for Multi-Turn Agentic Tasks
- MetaSyn: A Benchmark for LLM Agents on Meta-Analysis Articles from Nature Portfolio
- LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent
- Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
- Language-Coupled Reinforcement Learning for Multilingual Retrieval-Augmented Generation
- One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents
- AgentMath: Empowering Mathematical Reasoning for Large Language Models via Tool-Augmented Agent
- Laser: Governing Long-Horizon Agentic Search via Structured Protocol and Context Register
- Memory-T1: Reinforcement Learning for Temporal Reasoning in Multi-session Agents
- GenEnv: Difficulty-Aligned Co-Evolution Between LLM Agents and Environment Simulators
- Training LLMs with LogicReward for Faithful and Rigorous Reasoning
- Reinforcement Learning for Self-Improving Agent with Skill Library
- Turn-PPO: Turn-Level Advantage Estimation with PPO for Improved Multi-Turn RL in Agentic LLMs
- AdaSearch: Balancing Parametric Knowledge and Search in Large Language Models via Reinforcement Learning
- Needle in the Web: A Benchmark for Retrieving Targeted Web Pages in the Wild
- ToolForge: A Data Synthesis Pipeline for Multi-Hop Search without Real-World APIs
- An Open and Reproducible Deep Research Agent for Long-Form Question Answering
- ScholarGym: Benchmarking Large Language Model Capabilities in the Information-Gathering Stage of Deep Research
- MMhops-R1: Multimodal Multi-hop Reasoning
- AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning
- SpeakRL: Synergizing Reasoning, Speaking, and Acting in Language Models with Reinforcement Learning
- CoDA: A Context-Decoupled Hierarchical Agent with Reinforcement Learning
- PathFinder: MCTS and LLM Feedback-based Path Selection for Multi-Hop Question Answering
- KBQA-R1: Reinforcing Large Language Models for Knowledge Base Question Answering
- RouteRAG: Efficient Retrieval-Augmented Generation from Text and Graph via Reinforcement Learning
- LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services
- Enhancing Agentic RL with Progressive Reward Shaping and Value-based Sampling Policy Optimization
- An Index-based Approach for Efficient and Effective Web Content Extraction
- LightSearcher: Efficient DeepSearch via Experiential Memory
- VG-Refiner: Towards Tool-Refined Referring Grounded Reasoning via Agentic Reinforcement Learning
- Training Multi-Image Vision Agents via End2End Reinforcement Learning
- Distilling Temporal Search and Reasoning: Evolving LLMs for Future Prediction via Harness-Assisted Efficient Data Synthesis
- Thinking with Programming Vision: Towards a Unified View for Thinking with Images
- GTM: Simulating the World of Tools for AI Agents
- CARL: Criticality-Aware Agentic Reinforcement Learning
- On Group Relative Policy Optimization Collapse in Agent Search: The Lazy Likelihood-Displacement
- SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL
- Guided Self-Evolving LLMs with Minimal Human Supervision
- GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning
- Skywork-R1V4: Toward Agentic Multimodal Intelligence through Interleaved Thinking with Images and DeepResearch
- Agentic Policy Optimization via Instruction-Policy Co-Evolution
- CoSineVerifier: Tool-Augmented Answer Verification for Computation-Oriented Scientific Questions
- ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering
- ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration
- Reducing Latency of LLM Search Agent via Speculation-based Algorithm-System Co-Design
- Qwen3-VL Technical Report
- Thinking in 360°: Humanoid Visual Search in the Wild
- Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning
- Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization
- Latent Collaboration in Multi-Agent Systems
- CLaRa: Bridging Retrieval and Generation with Continuous Latent Reasoning
- Budget-Aware Tool-Use Enables Effective Agent Scaling
- Goal-Directed Search Outperforms Goal-Agnostic Memory Compression in Long-Context Memory Tasks
- SkyRL-Agent: Efficient RL Training for Multi-turn LLM Agent
- Agent0: Unleashing Self-Evolving Agents from Zero Data via Tool-Integrated Reasoning
- VisPlay: Self-Evolving Vision-Language Models from Images
- Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization
- Agent-R1: A Unified and Modular Framework for Agentic Reinforcement Learning
- Grounded by Experience: Generative Healthcare Prediction Augmented with Hierarchical Agentic Retrieval
- Explore More, Learn Better: Parallel MLLM Embeddings under Mutual Information Minimization
- LiveSearchBench: An Automatically Constructed Benchmark for Retrieval and Reasoning over Dynamic Knowledge
- Tailored Primitive Initialization is the Secret Key to Reinforcement Learning
- Constructing and Interpreting Digital Twin Representations for Visual Reasoning via Reinforcement Learning
- CriticSearch: Fine-Grained Credit Assignment for Search Agents via a Retrospective Critic
- Beyond ReAct: A Planner-Centric Framework for Complex Tool-Augmented LLM Reasoning
- AdaptFly: Prompt-Guided Adaptation of Foundation Models for Low-Altitude UAV Networks
- Thinking Forward and Backward: Multi-Objective Reinforcement Learning for Retrieval-Augmented Reasoning
- DEEPAMBIGQA: Ambiguous Multi-hop Questions for Benchmarking LLM Answer Completeness
- DeepSpecs: Expert-Level Questions Answering in 5G
- Thinker: Training LLMs in Hierarchical Thinking for Deep Search via Multi-Turn Interaction
- MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation
- From Experience to Strategy: Empowering LLM Agents with Trainable Graph Memory
- From Exploration to Exploitation: A Two-Stage Entropy RLVR Approach for Noise-Tolerant MLLM Training
- Q-RAG: Long Context Multi-step Retrieval via Value-based Embedder Training
- IterResearch: Rethinking Long-Horizon Agents via Markovian State Reconstruction
- Rethinking Retrieval-Augmented Generation for Medicine: A Large-Scale, Systematic Expert Evaluation and Practical Insights
- Think Before You Retrieve: Learning Test-Time Adaptive Search with Small Language Models
- DeepEyesV2: Toward Agentic Multimodal Model
- MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning
- Controlling Performance and Budget of a Centralized Multi-agent LLM System with Reinforcement Learning
- Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation
- Training Proactive and Personalized LLM Agents
- Beyond Single Embeddings: Capturing Diverse Targets with Multi-Query Retrieval
- Tool Zero: Training Tool-Augmented LLMs via Pure RL from Scratch
- Interaction as Intelligence Part II: Asynchronous Human-Agent Rollout for Long-Horizon Task Training
- Interact-RAG: Reason and Interact with the Corpus, Beyond Black-Box Retrieval
- ToolScope: An Agentic Framework for Vision-Guided and Long-Horizon Tool Use
- SIGMA: Search-Augmented On-Demand Knowledge Integration for Agentic Mathematical Reasoning
- MARAG-R1: Beyond Single Retriever via Reinforcement-Learned Multi-Tool Agentic Retrieval
- InfoFlow: Reinforcing Search Agent Via Reward Density Optimization
- Empowering RepoQA-Agent based on Reinforcement Learning Driven by Monte-carlo Tree Search
- Graph-Enhanced Policy Optimization in LLM Agent Training
- Towards Global Retrieval Augmented Generation: A Benchmark for Corpus-Level Reasoning
- One Model to Critique Them All: Rewarding Agentic Tool-Use via Efficient Reasoning
- PORTool: Tool-Use LLM Training with Rewarded Tree
- Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning
- Completion ≠ Collaboration: Scaling Collaborative Effort with Agents
- ALDEN: Reinforcement Learning for Active Navigation and Evidence Gathering in Long Documents
- MTIR-SQL: Multi-turn Tool-Integrated Reasoning Reinforcement Learning for Text-to-SQL
- CRMWeaver: Building Powerful Business Agent via Agentic RL and Shared Memories
- GAP: Graph-Based Agent Planning with Parallel Tool Use and Reinforcement Learning
- BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms
- When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation
- RWGBench: Evaluating Scholarly Positioning in Related Work Generation
- VerIF: Verification Engineering for Reinforcement Learning in Instruction Following
- AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery
- Evaluation of Agents under Simulated AI Marketplace Dynamics
- MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue
- Can Knowledge-Graph-based Retrieval Augmented Generation Really Retrieve What You Need?
- Sharpness-Guided Group Relative Policy Optimization via Probability Shaping
- Model-Document Protocol for AI Search
- KnowCoder-A1: Incentivizing Agentic Reasoning Capability with Outcome Supervision for KBQA
- AgentFrontier: Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data Synthesis
- Repurposing Synthetic Data for Fine-grained Search Agent Supervision
- OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning
- What Limits Agentic Systems Efficiency?
- LimRank: Less is More for Reasoning-Intensive Information Reranking
- On the Faithfulness of Visual Thinking: Measurement and Enhancement
- Think before Recommendation: Autonomous Reasoning-enhanced Recommender
- Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
- MGFRec: Towards Reinforced Reasoning Recommendation with Multiple Groundings and Feedback
- Rethinking the Design of Reinforcement Learning-Based Deep Research Agents
- Plan Then Retrieve: Reinforcement Learning-Guided Complex Reasoning over Knowledge Graphs
- GlobalRAG: Enhancing Global Reasoning in Multi-hop Question Answering via Reinforcement Learning
- Mixture-of-Minds: Multi-Agent Reinforcement Learning for Table Understanding
- ToolScope: Enhancing LLM Agent Tool Use through Tool Merging and Context-Aware Filtering
- SALT: Step-level Advantage Assignment for Long-horizon Agents via Trajectory Graph
- NeSyPr: Neurosymbolic Proceduralization For Efficient Embodied Reasoning
- Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search
- Search Self-play: Pushing the Frontier of Agent Capability without Supervision
- WebSeer: Training Deeper Search Agents through Reinforcement Learning with Self-Reflection
- Query Decomposition for RAG: Balancing Exploration-Exploitation
- Proactive Reasoning-with-Retrieval Framework for Medical Multimodal Large Language Models
- Agentic Reinforcement Learning for Search Misaligns Instruction-Tuning
- Rethinking On-policy Optimization for Query Augmentation
- SafeSearch: Do Not Trade Safety for Utility in LLM Search Agents
- Prompt-MII: Meta-Learning Instruction Induction for LLMs
- A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Evaluations, and Applications
- AutoGraph-R1: End-to-End Reinforcement Learning for Knowledge Graph Construction
- Cost-Aware Retrieval-Augmentation Reasoning Models with Adaptive Retrieval Depth
- MARSHAL: Incentivizing Multi-Agent Reasoning via Self-Play with Strategic LLMs
- EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle
- Structure-R1: Dynamically Leveraging Structural Knowledge in LLM Reasoning through Reinforcement Learning
- An Efficient Rubric-based Generative Verifier for Search-Augmented LLMs
- Internalizing World Models via Self-Play Finetuning for Agentic RL
- Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn LLM Agents
- Mapping Smarter, Not Harder: A Test-Time Reinforcement Learning Agent That Improves Without Labels or Model Updates
- A Hybrid, Knowledge-Guided Evolutionary Framework for Personalized Compiler Auto-Tuning
- Towards Agentic Self-Learning LLMs in Search Environment
- FinDeepResearch: Evaluating Deep Research Agents in Rigorous Financial Analysis
- ChatR1: Reinforcement Learning for Conversational Reasoning and Retrieval Augmented Question Answering
- Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation
- RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems
- LLM-guided Hierarchical Search for End-to-end Reasoning Intensive Retrieval
- DeepPlanner: Scaling Planning Capability for Deep Research Agents via Advantage Shaping
- DeepMMSearch-R1: Empowering Multimodal LLMs in Multimodal Web Search
- Memory as Action: Autonomous Context Curation for Long-Horizon Agentic Tasks
- Deep Research Brings Deeper Harm
- Demystifying Reinforcement Learning in Agentic Reasoning
- REGENT: Relevance-Guided Attention for Entity-Aware Multi-Vector Neural Re-Ranking
- A2FM: An Adaptive Agent Foundation Model for Tool-Aware Hybrid Reasoning
- Tree-based Dialogue Reinforced Policy Optimization for Red-Teaming Attacks
- Can Tool-Integrated Reinforcement Learning Generalize Across Diverse Domains?
- DyKnow-RAG: Dynamic Knowledge Utilization Reinforcement Framework for Noisy Retrieval-Augmented Generation in E-commerce Search Relevance
- DeepResearchGuard: Deep Research with Open-Domain Evaluation and Multi-Stage Guardrails for Safety
- A Survey on Agentic Multimodal Large Language Models
- Scaling Long-Horizon LLM Agent via Context-Folding
- MTSQL-R1: Towards Long-Horizon Multi-Turn Text-to-SQL via Agentic Training
- BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions
- GraphTracer: Graph-Guided Failure Tracing in LLM Agents for Robust Multi-Turn Deep Search
- RECON: Reasoning with Condensation for Efficient Retrieval-Augmented Generation
- Don't Just Fine-tune the Agent, Tune the Environment
- Beyond the limitation of a single query: Train your LLM for query expansion with Reinforcement Learning
- Veri-R1: Toward Precise and Faithful Claim Verification via Online Reinforcement Learning
- Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipe
- Can RL Improve Generalization of LLM Agents? An Empirical Study
- Autonomous Agents for Scientific Discovery: Orchestrating Scientists, Language, Code, and Physics
- How can we assess human-agent interactions? Case studies in software agent design
- LLP: LLM-based Product Pricing in E-commerce
- CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts
- DSPO: Stable and Efficient Policy Optimization for Agentic Search and Reasoning
- Agentic-KGR: Co-evolutionary Knowledge Graph Construction through Multi-Agent Reinforcement Learning
- GTAlign: Game-Theoretic Alignment of LLM Assistants for Social Welfare
- COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context
- Agent Learning via Early Experience
- Beyond Turn Limits: Training Deep Search Agents with Dynamic Context Window
- HiPRAG: Hierarchical Process Rewards for Efficient Agentic Retrieval Augmented Generation
- Understanding DeepResearch via Reports
- QAgent: A modular Search Agent with Interactive Query Understanding
- A2Search: Ambiguity-Aware Question Answering with Reinforcement Learning
- Training-Free Group Relative Policy Optimization
- Customer-R1: Personalized Simulation of Human Behaviors via RL-based LLM Agent in Online Shopping
- Search-R3: Unifying Reasoning and Embedding in Large Language Models
- Tool-Augmented Policy Optimization: Synergizing Reasoning and Adaptive Tool Use with Reinforcement Learning
- IoDResearch: Deep Research on Private Heterogeneous Data via the Internet of Data
- Scaling LLM Multi-turn RL with End-to-end Summarization-based Context Management
- AdvEvo-MARL: Shaping Internalized Safety through Adversarial Co-Evolution in Multi-Agent Reinforcement Learning
- Beneficial Reasoning Behaviors in Agentic Search and Effective Post-training to Obtain Them
- λ-GRPO: Unifying the GRPO Frameworks with Learnable Token Preferences
- TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics
- The Cognitive Bandwidth Bottleneck: Shifting Long-Horizon Agent from Planning with Actions to Planning with Schemas
- Off-Trajectory Reasoning: Can LLMs Collaborate on Reasoning Trajectory?
- Stratified GRPO: Handling Structural Heterogeneity in Reinforcement Learning of LLM Search Agents
- DecEx-RAG: Boosting Agentic Retrieval-Augmented Generation with Decision and Execution Optimization via Process Supervision
- Multi-Agent Tool-Integrated Policy Optimization
- TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use
- AlphaApollo: Orchestrating Foundation Models and Professional Tools into a Self-Evolving System for Deep Agentic Reasoning
- RLRF: Competitive Search Agent Design via Reinforcement Learning from Ranker Feedback
- AgentRL: Scaling Agentic Reinforcement Learning with a Multi-Turn, Multi-Task Framework
- TaskCraft: Automated Generation of Agentic Tasks
- Token Hidden Reward: Steering Exploration-Exploitation in Group Relative Deep Reinforcement Learning
- GEM: A Gym for Agentic LLMs
- MASH: Modeling Abstention via Selective Help-Seeking
- ReSeek: A Self-Correcting Framework for Search Agents with Instructive Rewards
- Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs
- Efficient and Transferable Agentic Knowledge Graph RAG via Reinforcement Learning
- MR2-Bench: Going Beyond Matching to Reasoning in Multimodal Retrieval
- Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners
- RE-Searcher: Robust Agentic Search with Goal-oriented Planning and Self-reflection
- RoRecomp: Enhancing Reasoning Efficiency via Rollout Response Recomposition in Reinforcement Learning
- TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
- TUMIX: Multi-Agent Test-Time Scaling with Tool-Use Mixture
- Planner-R1: Reward Shaping Enables Efficient Agentic RL with Smaller LLMs
- Hybrid Reward Normalization for Process-supervised Non-verifiable Agentic Tasks
- InfoAgent: Advancing Autonomous Information-Seeking Agents
- Flash-Searcher: Fast and Effective Web Agents via DAG-Based Parallel Execution
- Scaling Generalist Data-Analytic Agents
- Retro*: Optimizing LLMs for Reasoning-Intensive Document Retrieval
- AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play
- G-reasoner: Foundation Models for Unified Reasoning over Graph-structured Knowledge
- Bridging On-Device and Cloud LLMs for Collaborative Reasoning: A Unified Methodology for Local Routing and Post-Training
- Rethinking Reward Miscalibration of GRPO in Agentic RL
- Beyond the Exploration-Exploitation Trade-off: A Hidden State Approach for LLM Reasoning in RLVR
- SafeSearch: Automated Red-Teaming for the Safety of LLM-Based Search Agents
- PaperSearchQA: Learning to Search and Reason over Scientific Papers with RLVR
- Structured In-context Environment Scaling for Large Language Model Reasoning
- Toward Effective Tool-Integrated Reasoning via Self-Evolved Preference Learning
- Tagging the Thought: Unlocking Personalization Reasoning via Reinforcement Learning
- Look Back to Reason Forward: Revisitable Memory for Long-Context LLM Agents
- Language Models Can Learn from Verbal Feedback Without Scalar Rewards
- Variational Reasoning for Language Models
- Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning
- EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning
- Do LLM Agents Know How to Ground, Recover, and Assess? A Benchmark for Epistemic Competence in Information-Seeking Agents
- Think Right, Not More: Test-Time Scaling for Numerical Claim Verification
- GraphSearch: An Agentic Deep Searching Workflow for Graph Retrieval-Augmented Generation
- DeepTravel: An End-to-End Agentic Reinforcement Learning Framework for Autonomous Travel Planning Agents
- ResT: Reshaping Token-Level Policy Gradients for Tool-Use Large Language Models
- Tree Search for LLM Agent Reinforcement Learning
- SFT Doesn't Always Hurt General Capabilities: Revisiting Domain-Specific Fine-Tuning in LLMs
- Training Task Reasoning LLM Agents for Multi-turn Task Planning via Single-turn Reinforcement Learning
- RAR2: Retrieval-Augmented Medical Reasoning via Thought-Driven Retrieval
- UserRL: Training Interactive User-Centric Agent via Reinforcement Learning
- From Scoring to Acting: Outcome-Verified Comparative Self-Distillation for LLM Agents
- SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
- Harness-G: A Graph-Structured Harness for Search Agents
- VIG-RL: Learning to Search and Insert for Verified Image Grounding
- Group-Reflective Self-Distillation for Agentic Reinforcement Learning
- Contrastive Reinforced Policy Optimization via Privileged Self-Distillation
- Dr-DCI: Scaling Direct Corpus Interaction via Dynamic Workspace Expansion
- Agentic Reinforcement Learning with Implicit Step Rewards
- Pathways of Thoughts: Multi-Directional Thinking for Long-form Personalized Question Answering
- Cortex: Achieving Low-Latency, Cost-Efficient Remote Data Access For LLM via Semantic-Aware Knowledge Caching
- A State-Update Prompting Strategy for Efficient and Robust Multi-turn Dialogue
- Governing Automated Strategic Intelligence
- Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle
- ToolSample: Dual Dynamic Sampling Methods with Curriculum Learning for RL-based Tool Learning
- Improving Context Fidelity via Native Retrieval-Augmented Reasoning
- Omne-R1: Learning to Reason with Memory for Multi-hop Question Answering
- Tool-R1: Sample-Efficient Reinforcement Learning for Agentic Tool Use
- WebResearcher: Unleashing unbounded reasoning capability in Long-Horizon Agents
- ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization
- ToolRM: Outcome Reward Models for Tool-Calling Large Language Models
- DeepDive: Advancing Deep Search Agents with Knowledge Graphs and Multi-Turn RL
- Towards Secure and Explainable Smart Contract Generation with Security-Aware Group Relative Policy Optimization
- CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
- Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents
- AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
- Reinforcement Learning Foundations for Deep Research Systems: A Survey
- SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
- What-If Analysis of Large Language Models: Explore the Game World Using Proactive Thinking
- Fishing for Answers: Exploring One-shot vs. Iterative Retrieval Strategies for Retrieval Augmented Generation
- SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
- OwkinZero: Accelerating Biological Discovery with AI
- Batch Query Processing and Optimization for Agentic Workflows
- HF-RAG: Hierarchical Fusion-based RAG with Multiple Sources and Rankers
- Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
- VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
- Multimodal Iterative RAG for Knowledge-Intensive Visual Question Answering
- EviNote-RAG: Enhancing RAG Models via Answer-Supportive Evidence Notes
- Open Data Synthesis For Deep Research
- MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents
- Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models
- The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management
- PVPO: Pre-Estimated Value-Based Policy Optimization for Agentic Reasoning
- AI-SearchPlanner: Modular Agentic Search via Pareto-Optimal Multi-Objective Reinforcement Learning
- SurGE: A Benchmark and Evaluation Framework for Scientific Survey Generation
- Can Compact Language Models Search Like Agents? Distillation-Guided Policy Optimization for Preserving Agentic RAG Capabilities
- Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning
- Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning
- Improving Low-Resource Translation with Dictionary-Guided Fine-Tuning and RL: A Spanish-to-Wayuunaiki Study
- Understanding Tool-Integrated Reasoning
- Hybrid Deep Searcher: Integrating Parallel and Sequential Search Reasoning
- Test-time Corpus Feedback: From Retrieval to RAG
- MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use
- Harnessing Rule-Based Reinforcement Learning for Enhanced Grammatical Error Correction
- MedResearcher-R1: Expert-Level Medical Deep Researcher via A Knowledge-Informed Trajectory Synthesis Framework
- Revisiting RAG Ensemble: A Theoretical and Mechanistic Analysis of Multi-RAG System Collaboration
- Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward
- OS-R1: Agentic Operating System Kernel Tuning with Reinforcement Learning
- ToolACE-MT: Non-Autoregressive Generation for Agentic Multi-Turn Interaction
- Towards Open-Ended Emotional Support Conversations in LLMs via Reinforcement Learning with Future-Oriented Rewards
- Deep Research: A Survey of Autonomous Research Agents
- Reinforcement Learning with Rubric Anchors
- Simple o3: Towards Interleaved Vision-Language Reasoning
- SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling
- SSRL: Self-Search Reinforcement Learning
- MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
- ReviewRL: Towards Automated Scientific Review with RL
- ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning
- Reducing Cognitive Overhead in Tool Use via Multi-Small-Agent Reinforcement Learning
- BrowseMaster: Towards Scalable Web Browsing via Tool-Augmented Programmatic Agent Pair
- MuaLLM: A Multimodal Large Language Model Agent for Circuit Design Assistance with Hybrid Contextual Retrieval-Augmented Generation
- Audio-Thinker: Guiding Audio Language Model When and How to Think via Reinforcement Learning
- TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding
- Careful Queries, Credible Results: Teaching RAG Models Advanced Web Search Tools with Reinforcement Learning
- WideSearch: Benchmarking Agentic Broad Info-Seeking
- HierSearch: A Hierarchical Enterprise Deep Search Framework Integrating Local and Web Searches
- Beyond Ten Turns: Unlocking Long-Horizon Agentic Search with Large-Scale Asynchronous RL
- M2IO-R1: An Efficient RL-Enhanced Reasoning Framework for Multimodal Retrieval Augmented Multimodal Generation
- UR2: Unify RAG and Reasoning through Reinforcement Learning
- BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent
- Multi-module GRPO: Composing Policy Gradients and Prompt Optimization for Language Model Programs
- Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
- GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
- PAIRS: Parametric-Verified Adaptive Information Retrieval and Selection for Efficient RAG
- VeriGUI: Verifiable Long-Chain GUI Dataset
- Agent Lightning: Train ANY AI Agents with Reinforcement Learning
- RAVine: Reality-Aligned Evaluation for Agentic Search
- Tool-integrated Reinforcement Learning for Repo Deep Search
- Patho-AgenticRAG: Towards Multimodal Agentic Retrieval-Augmented Generation for Pathology VLMs via Reinforcement Learning
- A Survey on AgentOps: Categorization, Challenges, and Future Directions
- Don't Overthink It: A Survey of Efficient R1-style Large Reasoning Models
- A Survey of LLM-based Deep Search Agents: Paradigm, Optimization, Evaluation, and Challenges
- RepoForge: Training a SOTA Fast-thinking SWE Agent with an End-to-End Data Curation Pipeline Synergizing SFT and RL at Scale
- Prompting Large Language Models with Partial Knowledge for Answering Questions with Unseen Entities
- VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
- PilotRL: Training Language Model Agents via Global Planning-Guided Progressive Reinforcement Learning
- MetaAgent: Toward Self-Evolving Agent via Tool Meta-Learning
- MAO-ARAG: Multi-Agent Orchestration for Adaptive Retrieval-Augmented Generation
- GraphRAG-R1: Graph Retrieval-Augmented Generation with Process-Constrained Reinforcement Learning
- Annotation-Free Reinforcement Learning Query Rewriting via Verifiable Search Reward
- Interaction as Intelligence: Deep Research With Human-AI Partnership
- From Sufficiency to Reflection: Reinforcement-Guided Thinking Quality in Retrieval-Augmented Reasoning for LLMs
- Learning to Extract Rational Evidence via Reinforcement Learning for Retrieval-Augmented Generation
- A Comprehensive Taxonomy of Negation for NLP and Neural Retrievers
- UserBench: An Interactive Gym Environment for User-Centric Agents
- Graph-R1: Towards Agentic GraphRAG Framework via End-to-end Reinforcement Learning
- RAG in the Wild: On the (In)effectiveness of LLMs with Mixture-of-Knowledge Retrieval Augmentation
- Agentic Reinforced Policy Optimization
- Does More Inference-Time Compute Really Help Robustness?
- WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization
- AgentFly: Extensible and Scalable Reinforcement Learning for LM Agents
- PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation
- MemGraphRAG: Memory-based Multi-Agent System for Graph Retrieval-Augmented Generation
- FutureSim: Replaying World Events to Evaluate Adaptive Agents
- Shop-R1: Rewarding LLMs to Simulate Human Behavior in Online Shopping via Reinforcement Learning
- DynaSearcher: Dynamic Knowledge Graph Augmented Search Agent via Multi-Reward Reinforcement Learning
- Deliberative Searcher: Improving LLM Reliability via Reinforcement Learning with constraints
- A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning
- ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry
- Xiangqi-R1: Enhancing Spatial Strategic Reasoning in LLMs for Chinese Chess via Reinforcement Learning
- Router-R1: Teaching LLMs Multi-Round Routing and Aggregation via Reinforcement Learning
- A Survey on Large Language Models for Mathematical Reasoning
- VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across Domains
- Am I on the Right Track? What Can Predicted Query Performance Tell Us about the Search Behaviour of Agentic RAG
- Towards Agentic RAG with Deep Reasoning: A Survey of RAG-Reasoning Systems in LLMs
- Leanabell-Prover-V2: Verifier-integrated Reasoning for Formal Theorem Proving via Reinforcement Learning
- Infinite Video Understanding
- RAISE: Enhancing Scientific Reasoning in LLMs via Step-by-Step Retrieval
- Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember
- FrugalRAG: Learning to retrieve and reason for multi-hop QA
- Agentic-R1: Distilled Dual-Strategy Reasoning
- Think2Go: Generative Next POI Recommendation with LLM Reasoning
- Deep Research Comparator: A Platform For Fine-grained Human Annotations of Deep Research Agents
- ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL
- Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution
- R1-RE: Cross-Domain Relation Extraction with RLVR
- SciMaster: Towards General-Purpose Scientific AI Agents, Part I. X-Master as Foundation: Can We Lead on Humanity's Last Exam?
- TAPR: Enhancing LLM Performance with a Task-Aware Prompt Rewriter
- ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning
- HiRA: A Hierarchical Reasoning Framework for Decoupled Planning and Execution in Deep Search
- WebSailor: Navigating Super-human Reasoning for Web Agent
- MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent
- RLVER: Reinforcement Learning with Verifiable Emotion Rewards for Empathetic Agents
- Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning
- Frustratingly Simple Retrieval Improves Challenging, Reasoning-Intensive Benchmarks
- Towards Robustness: A Critique of Current Vector Database Assessments
- Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints
- L0: Reinforcement Learning to Become General Agents
- RAG-R1: Incentivizing the Search and Reasoning Capabilities of LLMs through Multi-query Parallelism
- Real-Time Progress Prediction in Reasoning Language Models
- Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning
- SEEA-R1: Tree-Structured Reinforcement Fine-Tuning for Self-Evolving Embodied Agents
- Coordinating Search-Informed Reasoning and Reasoning-Guided Search in Claim Verification
- DeepVideo-R1: Video Reinforcement Fine-Tuning via Difficulty-aware Regressive GRPO
- MMSearch-R1: Incentivizing LMMs to Search
- π-CoT: Prolog-Initialized Chain-of-Thought Prompting for Multi-Hop Question-Answering
- ReCode: Updating Code API Knowledge with Reinforcement Learning
- KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality
- CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning
- From Web Search towards Agentic Deep Research: Incentivizing Search with Reasoning Agents
- Deep Research Agents: A Systematic Examination And Roadmap
- Graphs Meet AI Agents: Taxonomy, Progress, and Future Opportunities
- JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent
- KAG-Thinker: Interactive Thinking and Deep Reasoning in LLMs via Knowledge-Augmented Generation
- From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation
- From Passive to Active Reasoning: Can Large Language Models Ask the Right Questions under Incomplete Information?
- MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents
- A Vision for Geo-Temporal Deep Research Systems: Towards Comprehensive, Transparent, and Reproducible Geo-Temporal Information Synthesis
- Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
- ReVeal: Self-Evolving Code Agents via Reliable Self-Verification
- KnowCoder-V2: Deep Knowledge Analysis
- Conversational Search: From Fundamentals to Frontiers in the LLM Era
- CIIR@LiveRAG 2025: Optimizing Multi-Agent Retrieval Augmented Generation through Self-Training
- AutoMind: Adaptive Knowledgeable Agent for Automated Data Science
- Reasoning RAG via System 1 or System 2: A Survey on Reasoning Agentic Retrieval-Augmented Generation for Industry Challenges
- Reinforcement learning fine-tuning of language model for instruction following and math reasoning
- OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation
- Resisting Contextual Interference in RAG via Parametric-Knowledge Reinforcement
Discussions
- Search-R1: Training LLMs to Reason and Leverage Search Engines with RL [hn, 101 points, 12 comments]
- Search-R1: Training LLMs to Reason and Leverage Search Engines with RL [bsky, 1 points, 0 comments]
- Link: arxiv.org/pdf/2503.09516 [bsky, 0 points, 0 comments]
- UIUC & UMASS-Amherst team SEARCH-R1 extension to Deepseek, enhances LLM reasoning by integrating reinforcement learning with multi-turn real-time search, improving accuracy and flexibility in question [bsky, 0 points, 0 comments]
- Search-R1: Training LLMs to Reason and Leverage Search Engines with RL https://arxiv.org/abs/2503.09516 [bsky, 0 points, 0 comments]
- Search-R1: Training LLMs to Reason and Leverage Search Engines with RL https://arxiv.org/abs/2503.09516 [comments] [78 points] [bsky, 0 points, 0 comments]
- Paper: arxiv.org/abs/2503.09516 Code: github.com/PeterGriffin... [bsky, 0 points, 1 comments]
- ⚡ Hackernews Top story: Search-R1: Training LLMs to Reason and Leverage Search Engines with RL [bsky, 0 points, 0 comments]
- Search-R1: Training LLMs to Reason and Leverage Search Engines with RL (arxiv.org) Main Link | Discussion [bsky, 0 points, 0 comments]
- Search-R1: Training LLMs to Reason and Leverage Search Engines with RL Article URL: https://arxiv... https://arxiv.org/abs/2503.09516 Event Attributes [bsky, 0 points, 0 comments]
- Search-R1: Training LLMs to Reason and Leverage Search Engines with RL https://arxiv.org/abs/2503.09516 (https://news.ycombinator.com/item?id=43563265) [bsky, 0 points, 0 comments]
- https://arxiv.org/abs/2503.09516 この論文では、大規模言語モデル(LLM)が検索エンジンを効果的に利用できるようにするSearch-R1という手法を紹介しています。 Search-R1は、強化学習(RL)を通じて、LLMが自律的に検索クエリを生成する方法を学習します。 実験結果から、Search-R1は既存のベースラインモデルよりも大幅に性能が向上することが示されて [bsky, 0 points, 0 comments]
- Search-R1: Training LLMs to Reason and Leverage Search Engines with RL https:// arxiv.org/abs/2503.09516 # arxiv # llm # llms [mastodon, 0 points, 0 comments]
- Search-R1: Training LLMs to Reason and Leverage Search Engines with RL https://arxiv.org/abs/2503.09516 (https://news.ycombinator.com/item?id=43563265) [bsky, 0 points, 0 comments]
Related