WebSailor: Navigating Super-human Reasoning for Web Agent
2025/07/03 by Li, Kuan, Zhang, Zhongwang, Yin, Huifeng +16 · 80 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2507.02592
Abstract
Transcending human cognitive limitations represents a critical frontier in LLM training. Proprietary agentic systems like DeepResearch have demonstrated superhuman capabilities on extremely complex information-seeking benchmarks such as BrowseComp, a feat previously unattainable. We posit that their success hinges on a sophisticated reasoning pattern absent in open-source models: the ability to systematically reduce extreme uncertainty when navigating vast information landscapes. Based on this insight, we introduce WebSailor, a complete post-training methodology designed to instill this crucial capability. Our approach involves generating novel, high-uncertainty tasks through structured sampling and information obfuscation, RFT cold start, and an efficient agentic RL training algorithm, Duplicating Sampling Policy Optimization (DUPO). With this integrated pipeline, WebSailor significantly outperforms all opensource agents in complex information-seeking tasks, matching proprietary agents' performance and closing the capability gap.
Cited by
- Replay Failures as Successes: Sample-Efficient Reinforcement Learning for Instruction Following
- MindWatcher: Toward Smarter Multimodal Tool-Integrated Reasoning
- FoldAct: Efficient and Stable Context Folding for Long-Horizon Search Agents
- Hybrid Analysis for Secure MCP Tool Use in LLM Agents
- EcomBench: Towards Holistic Evaluation of Foundation Agents in E-commerce
- An Index-based Approach for Efficient and Effective Web Content Extraction
- Think in Parallel, Answer as One: Logit Averaging for Open-Ended Reasoning
- CuES: A Curiosity-driven and Environment-grounded Synthesis Framework for Agentic RL
- ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration
- PRInTS: Reward Modeling for Long-Horizon Information Seeking
- RhinoInsight: Improving Deep Research through Control Mechanisms for Model Behavior and Context
- Budget-Aware Tool-Use Enables Effective Agent Scaling
- SkyRL-Agent: Efficient RL Training for Multi-turn LLM Agent
- Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks
- MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling
- Environment Scaling for Interactive Agentic Experience Collection: A Survey
- IterResearch: Rethinking Long-Horizon Agents via Markovian State Reconstruction
- MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning
- Interact-RAG: Reason and Interact with the Corpus, Beyond Black-Box Retrieval
- InfoFlow: Reinforcing Search Agent Via Reward Density Optimization
- CRMWeaver: Building Powerful Business Agent via Agentic RL and Shared Memories
- AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery
- Tongyi DeepResearch Technical Report
- AgentFold: Long-Horizon Web Agents with Proactive Context Management
- AgentFrontier: Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data Synthesis
- Repurposing Synthetic Data for Fine-grained Search Agent Supervision
- Co-Sight: Enhancing LLM-Based Agents via Conflict-Aware Meta-Verification and Trustworthy Reasoning with Structured Facts
- Rethinking the Design of Reinforcement Learning-Based Deep Research Agents
- DeepWideSearch: Benchmarking Depth and Width in Agentic Information Seeking
- Search Self-play: Pushing the Frontier of Agent Capability without Supervision
- Enterprise Deep Research: Steerable Multi-Agent Deep Research for Enterprise Analytics
- A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Evaluations, and Applications
- Explore to Evolve: Scaling Evolved Aggregation Logic via Proactive Online Exploration for Deep Research Agents
- Synthesizing Agentic Data for Web Agents with Progressive Difficulty Enhancement Mechanisms
- Demystifying Reinforcement Learning in Agentic Reasoning
- A2FM: An Adaptive Agent Foundation Model for Tool-Aware Hybrid Reasoning
- A Survey on Agentic Multimodal Large Language Models
- BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions
- Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipe
- Beneficial Reasoning Behaviors in Agentic Search and Effective Post-training to Obtain Them
- Pushing Test-Time Scaling Limits of Deep Search with Asymmetric Verification
- Multi-Agent Tool-Integrated Policy Optimization
- InfoAgent: Advancing Autonomous Information-Seeking Agents
- Scaling Generalist Data-Analytic Agents
- Toward Effective Tool-Integrated Reasoning via Self-Evolved Preference Learning
- Do LLM Agents Know How to Ground, Recover, and Assess? A Benchmark for Epistemic Competence in Information-Seeking Agents
- DeepTravel: An End-to-End Agentic Reinforcement Learning Framework for Autonomous Travel Planning Agents
- Automatic Red Teaming LLM-based Agents with Model Context Protocol Tools
- PromptCoT 2.0: Scaling Prompt Synthesis for Large Language Model Reasoning
- TAPO: Transition-Aware Policy Optimization for LLM Agents
- Contrastive Reinforced Policy Optimization via Privileged Self-Distillation
- Generalizable End-to-End Tool-Use RL with Synthetic CodeGym
- Towards General Agentic Intelligence via Environment Scaling
- WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement Learning
- WebWeaver: Structuring Web-Scale Evidence with Dynamic Outlines for Open-Ended Deep Research
- WebResearcher: Unleashing unbounded reasoning capability in Long-Horizon Agents
- ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization
- Scaling Agents via Continual Pre-training
- Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search
- Reinforcement Learning Foundations for Deep Research Systems: A Survey
- WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents
- SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
- Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
- VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
- Open Data Synthesis For Deep Research
- AI-SearchPlanner: Modular Agentic Search via Pareto-Optimal Multi-Objective Reinforcement Learning
- Understanding Tool-Integrated Reasoning
- MedResearcher-R1: Expert-Level Medical Deep Researcher via A Knowledge-Informed Trajectory Synthesis Framework
- Revisiting RAG Ensemble: A Theoretical and Mechanistic Analysis of Multi-RAG System Collaboration
- MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
- BrowseMaster: Towards Scalable Web Browsing via Tool-Augmented Programmatic Agent Pair
- WideSearch: Benchmarking Agentic Broad Info-Seeking
- Beyond Ten Turns: Unlocking Long-Horizon Agentic Search with Large-Scale Asynchronous RL
- BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent
- WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent
- SEA: Self-Evolution Agent with Step-wise Reward for Computer Use
- VeriGUI: Verifiable Long-Chain GUI Dataset
- A Survey of LLM-based Deep Search Agents: Paradigm, Optimization, Evaluation, and Challenges
- PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation
- QUEST: Training Frontier Deep Research Agents with Fully Synthetic Tasks
Related