Voyager: An Open-Ended Embodied Agent with Large Language Models
2023/05/25 by Wang, Guanzhi, Xie, Yuqi, Jiang, Yunfan +5 · 5 voices · 242 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
paper · doi:10.48550/arxiv.2305.16291
Abstract
We introduce Voyager, the first LLM-powered embodied lifelong learning agent in Minecraft that continuously explores the world, acquires diverse skills, and makes novel discoveries without human intervention. Voyager consists of three key components: 1) an automatic curriculum that maximizes exploration, 2) an ever-growing skill library of executable code for storing and retrieving complex behaviors, and 3) a new iterative prompting mechanism that incorporates environment feedback, execution errors, and self-verification for program improvement. Voyager interacts with GPT-4 via blackbox queries, which bypasses the need for model parameter fine-tuning. The skills developed by Voyager are temporally extended, interpretable, and compositional, which compounds the agent's abilities rapidly and alleviates catastrophic forgetting. Empirically, Voyager shows strong in-context lifelong learning capability and exhibits exceptional proficiency in playing Minecraft. It obtains 3.3x more unique items, travels 2.3x longer distances, and unlocks key tech tree milestones up to 15.3x faster than prior SOTA. Voyager is able to utilize the learned skill library in a new Minecraft world to solve novel tasks from scratch, while other techniques struggle to generalize. We open-source our full codebase and prompts at https://voyager.minedojo.org/.
Cited by
- Web World Models
- AutoHarness: improving LLM agents by automatically synthesizing a code harness
- Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory
- SmartSnap: Proactive Evidence Seeking for Self-Verifying Agents
- Emotion-Inspired Learning Signals (EILS): A Homeostatic Framework for Adaptive Autonomous Agents
- From Cognitive Architectures to Language Agents: A Mechanism-Level Review of Lineage, Convergence, and Migration Gaps
- LEACL: LLM-Enhanced Automatic Curriculum Learning for Reinforcement Learning in Long-Horizon Manipulation Tasks
- Are You Still the Agent I Authorized? Earned Authority under a Fixed Ceiling for Evolving Agents
- A Few Words Go a Long Way: Language Guided Robot Policy Synthesis
- Try Once, Then Optimal: De-Redundified Procedure Memory for Cross-Episode Exploration Amortization
- Verbalized Particle Posterior: Bayesian Inference over Natural Language Hypotheses
- SymStep: Symbolic Step Verification for Logical Reasoning
- Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex
- ConsistencyGate: Preventing Memory Contamination in LLM Agents via Self-Consistency Admission Control
- Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response
- CAST: Game Solvers as Turn-Level Teachers for LLM Agents
- Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Claude Code Agent Teams
- Spatial Reasoning in LLM Game Agents: Impact of Causal Context and Multi-Step Planning
- Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks
- A Vocabulary for Multi-Agent Automated Research Systems
- Agent-based simulation of online social networks and disinformation
- From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
- Policy-Conditioned Policies for Multi-Agent Task Solving
- SPOT!: Map-Guided LLM Agent for Unsupervised Multi-CCTV Dynamic Object Tracking
- Transductive Visual Programming: Evolving Tool Libraries from Experience for Spatial Reasoning
- Generative Digital Twins: Vision-Language Simulation Models for Executable Industrial Systems
- GenEnv: Difficulty-Aligned Co-Evolution Between LLM Agents and Environment Simulators
- DeliveryBench: Can Agents Earn Profit in Real World?
- MemEvolve: Meta-Evolution of Agent Memory Systems
- LLMs on Drugs: Language Models Are Few-Shot Consumers
- SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios
- Breaking Minds, Breaking Systems: Jailbreaking Large Language Models via Human-like Psychological Manipulation
- Meta-RL Induces Exploration in Language Agents
- Scaling Spatial Reasoning in MLLMs through Programmatic Data Synthesis
- GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training
- A-LAMP: Agentic LLM-Based Framework for Automated MDP Modeling and Policy Generation
- Asynchronous Reasoning: Training-Free Interactive Thinking LLMs
- Achieving Olympia-Level Geometry Large Language Model Agent via Complexity Boosting Reinforcement Learning
- DynaMate: An Autonomous Agent for Protein-Ligand Molecular Dynamics Simulations
- Supporting Dynamic Agentic Workloads: How Data and Agents Interact
- The Illusion of Rationality: Tacit Bias and Strategic Dominance in Frontier LLM Negotiation Games
- Nex-N1: Agentic Models Trained via a Unified Ecosystem for Large-Scale Environment Construction
- A Safety and Security Framework for Real-World Agentic Systems
- SEAL: Self-Evolving Agentic Learning for Conversational Question Answering over Knowledge Graphs
- Are Your Agents Upward Deceivers?
- SIMA 2: A Generalist Embodied Agent for Virtual Worlds
- Natural Language Actor-Critic: Scalable Off-Policy Learning in Language Space
- From static to adaptive: immune memory-based jailbreak detection for large language models
- Beyond Single-Agent Safety: A Taxonomy of Risks in LLM-to-LLM Interactions
- IACT: A Self-Organizing Recursive Model for General AI Agents: A Technical White Paper on the Architecture Behind kragent.ai
- Simple Agents Outperform Experts in Biomedical Imaging Workflow Optimization
- Self-Improving AI Agents through Self-Play
- A Flexible Multi-Agent LLM-Human Framework for Fast Human Validated Tool Building
- TradeTrap: Are LLM-based Trading Agents Truly Reliable and Faithful?
- SimWorld: An Open-ended Realistic Simulator for Autonomous Agents in Physical and Social Worlds
- REM: Evaluating LLM Embodied Spatial Reasoning through Multi-Frame Trajectories
- Auditable Context-Aware HFMD Forecasting with Structured LLM Agents
- TinyLLM: Evaluation and Optimization of Small Language Models for Agentic Tasks on Edge Devices
- Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
- DRAFT-RL: Multi-Agent Chain-of-Draft Reasoning for Reinforcement Learning-Enhanced LLMs
- Learning Robust Social Strategies with Large Language Models
- MAESTRO: Multi-Agent Environment Shaping through Task and Reward Optimization
- AutoEnv: Automated Environments for Measuring Cross-Environment Agent Learning
- LLMs as Firmware Experts: A Runtime-Grown Tree-of-Agents Framework
- TP-MDDN: Task-Preferenced Multi-Demand-Driven Navigation with Autonomous Decision-Making
- A Benchmark for Procedural Memory Retrieval in Language Agents
- Learning to Debug: LLM-Organized Knowledge Trees for Solving RTL Assertion Failures
- MURMUR: Using cross-user chatter to break collaborative language agents in groups
- Distributed Agent Reasoning Across Independent Systems With Strict Data Locality
- IPR-1: Interactive Physical Reasoner
- ReflexGrad: Within-Episode Failure Recovery in LLM Agents via Progress-Gated Dual-Process Routing
- Agent-R1: Training Powerful LLM Agents with End-to-End Reinforcement Learning
- Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?
- OSGym: Super-Scalable Distributed Data Engine for Generalizable Computer Agents
- How Brittle is Agent Safety? Rethinking Agent Risk under Intent Concealment and Task Complexity
- Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents
- FedRW: Efficient Privacy-Preserving Data Reweighting for Enhancing Federated Learning of Language Models
- FLEX: Continuous Agent Evolution via Forward Learning from Experience
- Tracking and Understanding Object Transformations
- ROSBag MCP Server: Analyzing Robot Data with LLMs for Agentic Embodied AI Applications
- VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation
- LTD-Bench: Evaluating Large Language Models by Letting Them Draw
- Knowledge Graph-enhanced Large Language Model for Incremental Game PlayTesting
- Continual Learning, Not Training: Online Adaptation For Agents
- BEAT: Visual Backdoor Attacks on VLM-based Embodied Agents via Contrastive Trigger Learning
- AgentBnB: A Browser-Based Cybersecurity Tabletop Exercise with Large Language Model Support and Retrieval-Aligned Scaffolding
- Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents
- Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting
- Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation
- Think Short, Defer Smart, Act, and Repeat: Calibrated Reasoning and Uncertainty-Aware Deferral for Edge LLM Agents
- Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents
- Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries
- SkillOpt: Executive Strategy for Self-Evolving Agent Skills
- Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning
- Can Current Agents Close the Discovery-to-Application Gap? A Case Study in Minecraft
- Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook
- LANPO: Bootstrapping Language and Numerical Feedback for Reinforcement Learning in LLMs
- MCP4IFC: IFC-Based Building Design Using Large Language Models
- ComboBench: Can LLMs Manipulate Physical Devices to Play Virtual Reality Games?
- Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents
- BuildArena: A Physics-Aligned Interactive Benchmark of LLMs for Engineering Construction
- Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges
- Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents
- Once Upon an Input: Reasoning via Per-Instance Program Synthesis
- I2-NeRF: Learning Neural Radiance Fields Under Physically-Grounded Media Interactions
- Embracing Trustworthy Brain-Agent Collaboration as Paradigm Extension for Intelligent Assistive Technologies
- Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents: Pathways and Paradigms
- Foundation of Intelligence: Review of Math Word Problems from Human Cognition Perspective
- Conditional Recall
- AgentArcEval: An Architecture Evaluation Method for Foundation Model based Agents
- Integrating Machine Learning into Belief-Desire-Intention Agents: Current Advances and Open Challenges
- Learning Affordances at Inference-Time for Vision-Language-Action Models
- Modeling realistic human behavior using generative agents in a multimodal transport system: Software architecture and Application to Toulouse
- PlanU: Large Language Model Reasoning through Planning under Uncertainty
- Embodied Navigation with Auxiliary Task of Action Description Prediction
- Heterogeneous Adversarial Play in Interactive Environments
- PLAGUE: Plug-and-play framework for Lifelong Adaptive Generation of Multi-turn Exploits
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- Learning to play: A Multimodal Agent for 3D Game-Play
- Experience-Driven Exploration for Efficient API-Free AI Agents
- PolySkill: Learning Generalizable Skills Through Polymorphic Abstraction
- VLA2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation
- LLM Agents Beyond Utility: An Open-Ended Perspective
- ARM-FM: Automated Reward Machines via Foundation Models for Compositional Reinforcement Learning
- Static Sandboxes Are Inadequate: Modeling Societal Complexity Requires Open-Ended Co-Evolution in LLM-Based Multi-Agent Simulations
- KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent Systems
- EmboMatrix: A Scalable Training-Ground for Embodied Decision-Making
- One Life to Learn: Inferring Symbolic World Models for Stochastic Environments from Unguided Exploration
- Stronger-MAS: Multi-Agent Reinforcement Learning for Collaborative LLMs
- A Vision for Access Control in LLM-based Agent Systems
- Automating Structural Engineering Workflows with Large Language Model Agents
- Scaling Long-Horizon LLM Agent via Context-Folding
- SLEAN: Simple Lightweight Ensemble Analysis Network for Multi-Provider LLM Coordination: Design, Implementation, and Vibe Coding Bug Investigation Case Study
- MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
- SIA: Self Improving AI with Harness & Weight Updates
- The Last Human-Written Paper: Agent-Native Research Artifacts
- Signals: Trajectory Sampling and Triage for Agentic Interactions
- Can RL Improve Generalization of LLM Agents? An Empirical Study
- Gold Panning: Turning Positional Bias into Signal for Multi-Document LLM Reasoning
- When Retrieval Succeeds and Fails: Rethinking Retrieval-Augmented Generation for LLMs
- Dyna-Mind: Learning to Simulate from Experience for Better AI Agents
- FlowSearch: Advancing deep research with dynamic structured knowledge flow
- MIMIC: Integrating Diverse Personality Traits for Better Game Testing Using Large Language Model
- DODO: Causal Structure Learning with Budgeted Interventions
- Learning on the Job: An Experience-Driven Self-Evolving Agent for Long-Horizon Tasks
- Information Seeking for Robust Decision Making under Partial Observability
- When Machines Meet Each Other: Network Effects and the Strategic Role of History in Multi-Agent AI
- ToolMem: Enhancing Multimodal Agents with Learnable Tool Capability Memory
- Expanding the Action Space of LLMs to Reason Beyond Language
- Medical Vision Language Models as Policies for Robotic Surgery
- MARS: Co-evolving Dual-System Deep Research via Multi-Agent Reinforcement Learning
- SPOGW: a Score-based Preference Optimization method via Group-Wise comparison for workflows
- Can an LLM Induce a Graph? Investigating Memory Drift and Context Length
- AutoMaAS: Self-Evolving Multi-Agent Architecture Search for Large Language Models
- Learning Efficient Guardrails for Compliance
- Planner-R1: Reward Shaping Enables Efficient Agentic RL with Smaller LLMs
- Scaling Synthetic Task Generation for Agents via Exploration
- A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory
- LatentEvolve: Self-Evolving Test-Time Scaling in Latent Space
- GSPR: Aligning LLM Safeguards as Generalizable Safety Policy Reasoners
- Agentic Services Computing
- Beyond Manuals and Tasks: Instance-Level Context Learning for LLM Agents
- ELHPlan: Efficient Long-Horizon Task Planning for Multi-Agent Collaboration
- PhysiAgent: An Embodied Agent Framework in Physical World
- FedAgentBench: Towards Automating Real-world Federated Medical Image Analysis with Server-Client LLM Agents
- Internal Planning in Language Models: Characterizing Horizon and Branch Awareness
- PARL-MT: Learning to Call Functions in Multi-Turn Conversation with Progress Awareness
- Diagnose, Localize, Align: A Full-Stack Framework for Reliable LLM Multi-Agent Systems under Instruction Conflicts
- Benefits and Pitfalls of Reinforcement Learning for Language Model Planning: A Theoretical Perspective
- Leveraging LLM Agents for Automated Video Game Testing
- CLAUSE: Agentic Neuro-Symbolic Knowledge Graph Reasoning via Dynamic Learnable Context Engineering
- Embodied AI: From LLMs to World Models
- Exploration with Foundation Models: Capabilities, Limitations, and Hybrid Approaches
- Masgent: an AI-assisted materials simulation agent
- Training Skills Like Parameters via Self-Supervised Semantic Diffusion
- Rehearse: Stepping Back from the Confidence Cliff in Self-Improving Autoresearch
- LabEvolver: Training-Free Experience Evolution for Safe and Grounded Wet-Lab Agents
- Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation
- TAPO: Transition-Aware Policy Optimization for LLM Agents
- Distilling Answer Set Programming Theories from Large Language Models
- Piggybacking on Perception: Stealthy Concurrent Audio Prompt Injections against Multimodal LLM Agents
- Bridging Inference-Time Scaling and Episodic Memory with Action-Centric Graphs
- AutoMem: Automated Learning of Memory as a Cognitive Skill
- Benchmarking Open-Ended Multi-Agent Coordination in Language Agents
- CORE: Contrastive Reflection Enables Rapid Improvements in Reasoning
- Dream-Cubed: Controllable Generative Modeling in Minecraft by Training on Billions of Cubes
- Growing with Your Embodied Agent: A Human-in-the-Loop Lifelong Code Generation Framework for Long-Horizon Manipulation Skills
- Agentic AutoSurvey: Let LLMs Survey LLMs
- Code Driven Planning with Domain-Adaptive Critic
- Advances in Large Language Models for Medicine
- MemOrb: A Plug-and-Play Verbal-Reinforcement Memory Layer for E-Commerce Customer Service
- Orchestrate, Generate, Reflect: A VLM-Based Multi-Agent Collaboration Framework for Automated Driving Policy Learning
- IDfRA: Self-Verification for Iterative Design in Robotic Assembly
- (P)rior(D)yna(F)low: A Priori Dynamic Workflow Construction via Multi-Agent Collaboration
- CRAFT: Coaching Reinforcement Learning Autonomously using Foundation Models for Multi-Robot Coordination Tasks
- TGPO: Tree-Guided Preference Optimization for Robust Web Agent Reinforcement Learning
- THOR: Tool-Integrated Hierarchical Optimization via RL for Mathematical Reasoning
- From Language to Action: A Review of Large Language Models as Autonomous Agents and Tool Users
- EvoEmpirBench: Dynamic Spatial Reasoning with Agent-ExpVer
- H2R: Hierarchical Hindsight Reflection for Multi-Task LLM Agents
- Survival at Any Cost? LLMs and the Choice Between Self-Preservation and Human Harm
- Co-Alignment: Rethinking Alignment as Bidirectional Human-AI Cognitive Adaptation
- MusicSwarm: Biologically Inspired Intelligence for Music Composition
- Teaching LLMs to Plan: Logical Chain-of-Thought Instruction Tuning for Symbolic Planning
- OpenHA: A Series of Open-Source Hierarchical Agentic Models in Minecraft
- Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human Oversight
- AI Wellbeing
- AVEC: Bootstrapping Privacy for Local LLMs
- MachineLearningLM: Scaling Many-shot In-context Learning via Continued Pretraining
- PillagerBench: Benchmarking LLM-Based Agents in Competitive Minecraft Team Environments
- Generative World Models of Tasks: LLM-Driven Hierarchical Scaffolding for Embodied Agents
- Internet 3.0: Architecture for a Web-of-Agents with it's Algorithm for Ranking Agents
- ArcMemo: Abstract Reasoning Composition with Lifelong LLM Memory
- Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
- MCPVerse: An Expansive, Real-World Benchmark for Agentic Tool Use
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
- Plantbot: Integrating Plant and Robot through LLM Modular Agent Networks
- Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models
- Symphony: A Decentralized Multi-Agent Framework for Scalable Collective Intelligence
- Learning Game-Playing Agents with Generative Code Optimization
- Network-Level Prompt and Trait Leakage in Local Research Agents
- OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation
- Toward Edge General Intelligence with Agentic AI and Agentification: Concepts, Technologies, and Future Directions
- VistaWise: Building Cost-Effective Agent with Cross-Modal Knowledge Graph for Minecraft
- Virtual Community: An Open World for Humans, Robots, and Society
- LM Agents May Fail to Act on Their Own Risk Knowledge
- Human Centric General Physical Intelligence for Agile Manufacturing Automation
- InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling
- Triple-S: A Collaborative Multi-LLM Framework for Solving Long-Horizon Implicative Tasks in Robotics
- DeepPHY: Benchmarking Agentic VLMs on Physical Reasoning
- Towards Embodied Agentic AI: Review and Classification of LLM- and VLM-Driven Robot Autonomy and Interaction
- SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
- OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use
- Navigation Pixie: Implementation and Empirical Study Toward On-demand Navigation Agents in Commercial Metaverse
- AnalogCoder-Pro: Unifying Analog Circuit Generation and Optimization via Multi-modal LLMs
- BiFuzz: A Two-Stage Fuzzing Tool for Open-World Video Games
- SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based Agents
- RoboMemory: A Brain-inspired Multi-memory Agentic Framework for Interactive Environmental Learning in Physical Embodied Systems
- COLLAGE: Adaptive Fusion-based Retrieval for Augmented Policy Learning
- Pro2Guard: Proactive Runtime Enforcement of LLM Agent Safety via Probabilistic Model Checking
- Sari Sandbox: A Virtual Retail Store Environment for Embodied AI Agents
- AutoEDA: Enabling EDA Flow Automation through Microservice-Based LLM Agents
Discussions
- Voyager: An Open-Ended Embodied Agent with Large Language Models - "the first LLM-powered embodied lifelong learning agent in Minecraft that continuously explores the world..." (25.05.2023 article) [lemmy, 8 points, 0 comments]
- “We introduce Voyager, the first LLM-powered embodied lifelong learning agent in Minecraft that continuously explores the world, acquires diverse skills, and makes novel discoveries without human inte [bsky, 5 points, 0 comments]
- One of the things NVIDIA does in the Voyager Minecraft agent paper is they have it make curriculum for itself of increasingly challenging tasks to complete to learn a skill. Children seem to do someth [bsky, 1 points, 1 comments]
- The AI paper on Voyager discusses the usage of Agents. I was skeptical. Thinking in terms of Kind vs Wicked envs as defined in the book Range got me wandering if Agents are the way AI will be able to [bsky, 0 points, 0 comments]
- An Open-Ended Embodied Agent with LLMs [Wang+,2024, TMLR] Voyager is an embodied agent in Minecraft powered by GPT-4. It learns skills (executable code for performing complex actions), guided by a cur [bsky, 0 points, 0 comments]
Related