WebGPT: Browser-assisted question-answering with human feedback
2021/12/17 by Reiichiro Nakano, Nakano, Reiichiro, Jacob Hilton +35 · 3 voices · 203 citations
Computer Science · #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2112.09332
openalex publication_date 2021/12/17 · openalex created_date 2022/05/05 · openalex updated_date 2026/07/28
Abstract
We fine-tune GPT-3 to answer long-form questions using a text-based web-browsing environment, which allows the model to search and navigate the web. By setting up the task so that it can be performed by humans, we are able to train models on the task using imitation learning, and then optimize answer quality with human feedback. To make human evaluation of factual accuracy easier, models must collect references while browsing in support of their answers. We train and evaluate our models on ELI5, a dataset of questions asked by Reddit users. Our best model is obtained by fine-tuning GPT-3 using behavior cloning, and then performing rejection sampling against a reward model trained to predict human preferences. This model's answers are preferred by humans 56% of the time to those of our human demonstrators, and 69% of the time to the highest-voted answer from Reddit.
Cited by
- WebCoach: Self-Evolving Web Agents with Cross-Session Memory Guidance
- Answer-Reconstruction Search Density: Measuring the Query and Source Work Compressed by Conversational Answers
- Do Data Agents Need Semantic Metadata? A Comparative Study in Agentic Data Retrieval
- ClawBench: Can AI Agents Complete Everyday Online Tasks?
- DeLIVeR: Decomposed Learning for Information-grounded Veracity Recognition via Reinforced Knowledge Graph Exploration
- The Behavioral Credibility Trilemma: When Calibrated Autonomy Becomes Impossible
- DRNOISE: Benchmarking Deep Research Agents in Misleading Evidence Environments
- Do Generative AI Assistants Respect robots.txt? Tracing Web Access Beyond Visible Answers
- RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources
- AnovaX: A Local, Multi-Agent Voice Assistant with LLM Planning, Typed Executors, and Adaptive Recovery
- Multi-Turn On-Policy Distillation with Prefix Replay
- SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration
- FormulaSPIN: Self-Play Fine-Tuning for Natural Language to Spreadsheet Formula Generation
- Hallucinations Undermine Trust; Metacognition is a Way Forward
- Skill-Contracted Agents for Evidence-Aware Materials Literature Analysis
- Calibrated Selective Fact-Checking via Evidence Chain Evaluation
- Learning to Orchestrate Agents in Natural Language with the Conductor
- Layer-0 Suppressors Ground Hallucination Inevitability: A Mechanistic Account of How Transformers Trade Factuality for Hedging
- K2-Think: A Parameter-Efficient Reasoning System
- Coral Protocol: Open Infrastructure Connecting The Internet of Agents
- It's LIT! Reliability-Optimized LLMs with Inspectable Tools
- SPIRAL: Symbolic LLM Planning via Grounded and Reflective Search
- Video-BrowseComp: Benchmarking Agentic Video Research on Open Web
- Agent2World: Learning to Generate Symbolic World Models via Adaptive Multi-Agent Feedback
- Unbiased Visual Reasoning with Controlled Visual Inputs
- ConsistencyGate: Preventing Memory Contamination in LLM Agents via Self-Consistency Admission Control
- Inverse RL Helps Align AI by Imitating Humans
- From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
- The Cartesian Cut in Agentic AI
- An Information Theoretic Perspective on Agentic System Design
- A Unified Definition of Hallucination, Or: It's the World Model, Stupid
- AgentMath: Empowering Mathematical Reasoning for Large Language Models via Tool-Augmented Agent
- AdaSearch: Balancing Parametric Knowledge and Search in Large Language Models via Reinforcement Learning
- Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential Game
- PDE-Agent: A toolchain-augmented multi-agent framework for PDE solving
- Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
- AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning
- VERAFI: Verified Agentic Financial Intelligence through Neurosymbolic Policy Generation
- Source Coverage and Citation Bias in LLM-based vs. Traditional Search Engines
- ValuePilot: A Two-Phase Framework for Value-Driven Decision-Making
- DART: Leveraging Multi-Agent Disagreement for Tool Recruitment in Multimodal Reasoning
- On the Limits of Test-Time Compute: Sequential Reward Filtering for Better Inference
- Overcoming State Inertia: Minimally Invasive Temporal Alignment for Evolving Contexts
- Process-Centric Analysis of Agentic Software Systems
- Evolving Paradigms in Task-Based Search and Learning: A Comparative Analysis of Traditional Search Engine with LLM-Enhanced Conversational Search System
- An Empirical Study on the Security Vulnerabilities of GPTs
- ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration
- VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering
- Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization
- CodeV: Code with Images for Faithful Visual Reasoning via Tool-Aware Policy Optimization
- SPINE: Token-Selective Test-Time Reinforcement Learning with Entropy-Band Regularization
- Goal-Directed Search Outperforms Goal-Agnostic Memory Compression in Long-Context Memory Tasks
- Fast LLM Post-training via Decoupled and Fastest-of-N Speculation
- Finetuning LLMs for Automatic Form Interaction on Web-Browser in Selenium Testing Framework
- Mobile-Agent-RAG: Driving Smart Multi-Agent Coordination with Contextual Knowledge Empowerment for Long-Horizon Mobile Automation
- Convergence and Stability Analysis of Self-Consuming Generative Models with Heterogeneous Human Curation
- AlphaCast: A Human Wisdom-LLM Intelligence Co-Reasoning Framework for Interactive Time Series Forecasting
- OSGym: Super-Scalable Distributed Data Engine for Generalizable Computer Agents
- From Experience to Strategy: Empowering LLM Agents with Trainable Graph Memory
- IterResearch: Rethinking Long-Horizon Agents via Markovian State Reconstruction
- Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs
- Inference-Time Personalized Alignment with a Few User Preference Queries
- Tool Zero: Training Tool-Augmented LLMs via Pure RL from Scratch
- A CPU-Centric Perspective on Agentic AI
- GraphChain: Large Language Models for Large-scale Graph Analysis via Tool Chaining
- PORTool: Tool-Use LLM Training with Rewarded Tree
- Decomposition-Enhanced Training for Post-Hoc Attributions In Language Models
- Truth-Aware Decoding: A Program-Logic Approach to Factual Language Generation
- What Does It Take to Detect an AI Agent? Minimal Feature Sets for Behavioral Detection under Browser Automation
- GroundEval: A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluation
- Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents
- A Survey of Data Agents: Emerging Paradigm or Overstated Hype?
- The Best of N Worlds: Aligning Reinforcement Learning with Best-of-N Sampling via max@k Optimisation
- Adaptive Blockwise Search: Inference-Time Alignment for Large Language Models
- Multi-Stakeholder Alignment in LLM-Powered Collaborative AI Systems: A Multi-Agent Framework for Intelligent Tutoring
- Beyond Reasoning Gains: Mitigating General Capabilities Forgetting in Large Reasoning Models
- Reducing the Probability of Undesirable Outputs in Language Models Using Probabilistic Inference
- PanicToCalm: A Proactive Counseling Agent for Panic Attacks
- Re-evaluating Minimum Bayes Risk Decoding for Automatic Speech Recognition
- Surfer 2: The Next Generation of Cross-Platform Computer Use Agents
- CUARewardBench: A Benchmark for Evaluating Reward Models on Computer-using Agent
- Crucible: Quantifying the Potential of Control Algorithms through LLM Agents
- Contextual Attention Modulation: Towards Efficient Multi-Task Adaptation in Large Language Models
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
- WEBSERV: A Full-Stack and RL-Ready Web Environment for Training Web Agents at Scale
- Natural Language Tools: A Natural Language Approach to Tool Calling In Large Language Agents
- NetMCP: Network-Aware Model Context Protocol Platform for LLM Capability Extension
- Grounding Long-Context Reasoning with Contextual Normalization for Retrieval-Augmented Generation
- Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanism
- Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
- On the Role of Preference Variance in Preference Optimization
- How Well Can Preference Optimization Generalize Under Noisy Feedback?
- Attacks by Content: Automated Fact-checking is an AI Security Issue
- SusBench: An Online Benchmark for Evaluating Dark Pattern Susceptibility of Computer-Use Agents
- A Survey on Agentic Multimodal Large Language Models
- PaperArena: An Evaluation Benchmark for Tool-Augmented Agentic Reasoning on Scientific Literature
- LLM×MapReduce-V3: Enabling Interactive In-Depth Survey Generation through a MCP-Driven Hierarchically Modular Agent System
- Known By Their Actions: Fingerprinting LLM Browser Agents via UI Traces
- Can RL Improve Generalization of LLM Agents? An Empirical Study
- Safety Game: Balancing Safe and Informative Conversations with Blackbox Agentic AI using LP Solvers
- Fundamentals of Building Autonomous LLM Agents
- MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning
- FlowSearch: Advancing deep research with dynamic structured knowledge flow
- Memory Retrieval and Consolidation in Large Language Models through Function Tokens
- CREST-Search: Comprehensive Red-teaming for Evaluating Safety Threats in Large Language Models Powered by Web Search
- xRouter: Training Cost-Aware LLMs Orchestration System via Reinforcement Learning
- CommandSans: Securing AI Agents with Surgical Precision Prompt Sanitization
- Prepared mind, fast response: A temporal decoupling framework for adaptive knowledge orchestration in open-domain dialogue
- Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts
- Tool-Augmented Policy Optimization: Synergizing Reasoning and Adaptive Tool Use with Reinforcement Learning
- MIRA: Towards Mitigating Reward Hacking in Inference-Time Alignment of T2I Diffusion Models
- Efficient Tree-Structured Deep Research with Adaptive Resource Allocation
- Exposing Citation Vulnerabilities in Generative Engines
- InfoMosaic-Bench: Evaluating Multi-Source Information Seeking in Tool-Augmented Agents
- Optimal Stopping vs Best-of-N for Inference Time Optimization
- Test-Time Scaling in Diffusion LLMs via Hidden Semi-Autoregressive Experts
- MARS: Co-evolving Dual-System Deep Research via Multi-Agent Reinforcement Learning
- AlphaApollo: Orchestrating Foundation Models and Professional Tools into a Self-Evolving System for Deep Agentic Reasoning
- Best of mini-N in-loop Sampling: A Contextual Quality Reward Model for Reliable and Efficient Best-of-N Sampling
- JEF-Hinter: Leveraging Offline Knowledge for Improving Web Agents Adaptation
- Best-of-Majority: Minimax-Optimal Strategy for Pass@k Inference Scaling
- AgenticRAG: Tool-Augmented Foundation Models for Zero-Shot Explainable Recommender Systems
- Demystifying deep search: a holistic evaluation with hint-free multi-hop questions and factorised metrics
- PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents
- Limited Preference Data? Learning Better Reward Model with Latent Space Synthesis
- A Framework for Studying AI Agent Behavior: Evidence from Consumer Choice Experiments
- Planner-R1: Reward Shaping Enables Efficient Agentic RL with Smaller LLMs
- Detecting and Fixing API Misuses of Data Science Libraries Using Large Language Models
- Structural Reward Model: Enhancing Interpretability, Efficiency, and Scalability in Reward Modeling
- Humanline: Online Alignment as Perceptual Loss
- Not Wrong, But Untrue: LLM Overconfidence in Document-Based Queries
- Mix-Ecom: Towards Mixed-Type E-Commerce Dialogues with Complex Domain Rules
- SafeSearch: Automated Red-Teaming for the Safety of LLM-Based Search Agents
- Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment
- Large-Scale Constraint Generation -- Can LLMs Parse Hundreds of Constraints?
- PARL-MT: Learning to Call Functions in Multi-Turn Conversation with Progress Awareness
- Dialogues with AI Reduce Beliefs in Misinformation but Build No Lasting Discernment Skills
- Fine-Grained Detection of Context-Grounded Hallucinations Using LLMs
- What Should I Cite? A RAG Benchmark for Academic Citation Prediction
- It's Not You, It's Clipping: A Soft Trust-Region via Probability Smoothing for LLM RL
- Hallucination-Resistant, Domain-Specific Research Assistant with Self-Evaluation and Vector-Grounded Retrieval
- ToolBrain: A Flexible Reinforcement Learning Framework for Agentic Tools
- SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
- DeepResearch Agent System
- Harness-G: A Graph-Structured Harness for Search Agents
- Dr-DCI: Scaling Direct Corpus Interaction via Dynamic Workspace Expansion
- Benchmarking Open-Ended Multi-Agent Coordination in Language Agents
- Reflect before Act: Proactive Error Correction in Language Models
- CoCoA: Confidence and Context-Aware Adaptive Decoding for Resolving Knowledge Conflicts in Large Language Models
- UIPro: Unleashing Superior Interaction Capability For GUI Agents
- Towards General Computer Control with Hierarchical Agents and Multi-Level Action Spaces
- Asking a Language Model for Diverse Responses
- SignalLLM: A General-Purpose LLM Agent Framework for Automated Signal Processing
- Governing Automated Strategic Intelligence
- SEFRQO: A Self-Evolving Fine-Tuned RAG-Based Query Optimizer
- RLinf: Flexible and Efficient Large-scale Reinforcement Learning via Macro-to-Micro Flow Transformation
- A Framework for Generating Artificial Datasets to Validate Absolute and Relative Position Concepts
- SIRAG: Towards Stable and Interpretable RAG with A Process-Supervised Multi-Agent Framework
- Environmental Injection Attacks against GUI Agents in Realistic Dynamic Environments
- DeepDive: Advancing Deep Search Agents with Knowledge Graphs and Multi-Turn RL
- WebSight: A Vision-First Architecture for Robust Web Agents
- Decoding Alignment: A Critical Survey of LLM Development Initiatives through Value-setting and Data-centric Lens
- Comp-X: On Defining an Interactive Learned Image Compression Paradigm With Expert-driven LLM Agent
- VehicleWorld: A Highly Integrated Multi-Device Environment for Intelligent Vehicle Interaction
- Icon2: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation
- Towards a Unified View of Large Language Model Post-Training
- Explainable Knowledge Graph Retrieval-Augmented Generation (KG-RAG) with KG-SMILE
- DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence
- Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
- EviNote-RAG: Enhancing RAG Models via Answer-Supportive Evidence Notes
- Can Compact Language Models Search Like Agents? Distillation-Guided Policy Optimization for Preserving Agentic RAG Capabilities
- ReSURE: Regularizing Supervision Unreliability for Multi-turn Dialogue Fine-tuning
- Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs
- From Bits to Boardrooms: A Cutting-Edge Multi-Agent LLM Framework for Business Excellence
- Foundational Design Principles and Patterns for Building Robust and Adaptive GenAI-Native Systems
- Better Language Model-Based Judging Reward Modeling through Scaling Comprehension Boundaries
- Select to Know: An Internal-External Knowledge Self-Selection Framework for Domain-Specific Question Answering
- Multimodal Data Storage and Retrieval for Embodied AI: A Survey
- A Functionality-Grounded Benchmark for Evaluating Web Agents in E-commerce Domains
- Deep Research: A Survey of Autonomous Research Agents
- Fast, Slow, and Tool-augmented Thinking for LLMs: A Review
- Thinking Inside the Mask: In-Place Prompting in Diffusion LLMs
- MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
- Improving and Evaluating Open Deep Research Agents
- OpenCUA: Open Foundations for Computer-Use Agents
- Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression
- HGMF: A Hierarchical Gaussian Mixture Framework for Scalable Tool Invocation within the Model Context Protocol
- Remote Sensing Image Intelligent Interpretation with the Language-Centered Perspective: Principles, Methods and Challenges
- Exploring a biocentric LLM-based assistant in environmental decision-making with more-than-human representation of the Tagus Estuary
- Can Large Models Fool the Eye? A New Turing Test for Biological Animation
- SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
- Multi-module GRPO: Composing Policy Gradients and Prompt Optimization for Language Model Programs
- Chain of Questions: Guiding Multimodal Curiosity in Language Models
- Large Language Model's Multi-Capability Alignment in Biomedical Domain
- ToolGrad: Efficient Tool-use Dataset Generation with Textual "Gradients"
- AttnTrace: Attention-based Context Traceback for Long-Context LLMs
- SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based Agents
- CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
- MetaAgent: Toward Self-Evolving Agent via Tool Meta-Learning
- BAR Conjecture: the Feasibility of Inference Budget-Constrained LLM Services with Authenticity and Reasoning
- Phi-Ground Tech Report: Advancing Perception in GUI Grounding
- Improving Generative Ad Text on Facebook using Reinforcement Learning
Discussions
Related