ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
2020/10/08 by Mohit Shridhar, Xingdi Yuan, Shridhar, Mohit +9 · 272 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Robotics (cs.RO) #Topic Modeling #cs.AI #cs.CL #cs.CV #cs.LG #cs.RO
paper · pdf · doi:10.48550/arxiv.2010.03768
ICLR 2021; Data, code, and videos are available at alfworld.github.io
openalex publication_date 2020/10/08 · arxiv created 2021/03/14 · arxiv updated 2021/03/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Given a simple request like Put a washed apple in the kitchen fridge, humans can reason in purely abstract terms by imagining action sequences and scoring their likelihood of success, prototypicality, and efficiency, all without moving a muscle. Once we see the kitchen in question, we can update our abstract plans to fit the scene. Embodied agents require the same abilities, but existing work does not yet provide the infrastructure necessary for both reasoning abstractly and executing concretely. We address this limitation by introducing ALFWorld, a simulator that enables agents to learn abstract, text based policies in TextWorld (Côté et al., 2018) and then execute goals from the ALFRED benchmark (Shridhar et al., 2020) in a rich visual environment. ALFWorld enables the creation of a new BUTLER agent whose abstract knowledge, learned in TextWorld, corresponds directly to concrete, visually grounded actions. In turn, as we demonstrate empirically, this fosters better agent generalization than training only in the visually grounded environment. BUTLER's simple, modular design factors the problem to allow researchers to focus on models for improving every piece of the pipeline (language understanding, planning, navigation, and visual scene understanding).
Citations
Cited by
- Unbiased Visual Reasoning with Controlled Visual Inputs
- Memex(RL): Scaling Long-Horizon LLM Agents via Indexed Experience Memory
- SymStep: Symbolic Step Verification for Logical Reasoning
- CAST: Game Solvers as Turn-Level Teachers for LLM Agents
- Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks
- ATOD: Annealed Turn-Aware On-Policy Distillation for Multi-Turn Agentic Tasks
- PatchWorld: Gradient-Free Optimization of Executable World Models for Agent Environments
- LookPlanGraph: Embodied Instruction Following Method with VLM Graph Augmentation
- TongSIM: A General Platform for Simulating Intelligent Machines
- GenEnv: Difficulty-Aligned Co-Evolution Between LLM Agents and Environment Simulators
- DeliveryBench: Can Agents Earn Profit in Real World?
- Learning Hierarchical Procedural Memory for LLM Agents through Bayesian Selection and Contrastive Refinement
- From Word to World: Can Large Language Models be Implicit Text-based World Models?
- Trust-Region Adaptive Policy Optimization
- Meta-RL Induces Exploration in Language Agents
- City Navigation in the Wild: Exploring Emergent Navigation from Web-Scale Knowledge in MLLMs
- Differentiable Evolutionary Reinforcement Learning
- GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training
- An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
- CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
- End-to-end PDDL Planning with Hardcoded and Dynamic Agents
- Reflecting with Two Voices: A Co-Adaptive Dual-Strategy Framework for LLM-Based Agent Decision Making
- ValuePilot: A Two-Phase Framework for Value-Driven Decision-Making
- Agent Skills Matter: Inferring Proprietary Skills from Execution Trajectories
- Mathematical Framing for Different Agent Strategies
- Reason-Plan-ReAct: A Reasoner-Planner Supervising a ReAct Executor for Complex Enterprise Tasks
- Inference-Time Distillation: Cost-Efficient Agents Without Fine-Tuning or Manual Prompt Engineering
- Transforming Monolithic Foundation Models into Embodied Multi-Agent Architectures for Human-Robot Collaboration
- Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
- AutoEnv: Automated Environments for Measuring Cross-Environment Agent Learning
- ReEXplore: Improving MLLMs for Embodied Exploration with Contextualized Retrospective Experience Replay
- A2Flow: Automating Agentic Workflow Generation via Self-Adaptive Abstraction Operators
- A Benchmark for Procedural Memory Retrieval in Language Agents
- AutoTool: Efficient Tool Selection for Large Language Model Agents
- ReflexGrad: Within-Episode Failure Recovery in LLM Agents via Progress-Gated Dual-Process Routing
- Agent-R1: A Unified and Modular Framework for Agentic Reinforcement Learning
- STEP: Success-Rate-Aware Trajectory-Efficient Policy Optimization
- Beyond ReAct: A Planner-Centric Framework for Complex Tool-Augmented LLM Reasoning
- Environment Scaling for Interactive Agentic Experience Collection: A Survey
- Balancing Multi-modal Sensor Learning via Multi-objective Optimization
- Scaling Agent Learning via Experience Synthesis
- Towards Understanding, Analyzing, and Optimizing Agentic AI Execution: A CPU-Centric Perspective
- Graph-Enhanced Policy Optimization in LLM Agent Training
- Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting
- SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution
- Think Short, Defer Smart, Act, and Repeat: Calibrated Reasoning and Uncertainty-Aware Deferral for Edge LLM Agents
- Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents
- SkillOpt: Executive Strategy for Self-Evolving Agent Skills
- Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning
- ComboBench: Can LLMs Manipulate Physical Devices to Play Virtual Reality Games?
- ReCAP: Recursive Context-Aware Reasoning and Planning for Large Language Model Agents
- ReCode: Unify Plan and Action for Universal Granularity Control
- DeepAgent: A General Reasoning Agent with Scalable Toolsets
- SALT: Step-level Advantage Assignment for Long-horizon Agents via Trajectory Graph
- NeSyPr: Neurosymbolic Proceduralization For Efficient Embodied Reasoning
- WebGraphEval: Multi-Turn Trajectory Evaluation for Web Agents using Graph Representation
- AgentChangeBench: A Multi-Dimensional Evaluation Framework for Goal-Shift Robustness in Conversational AI
- FineVision: Open Data Is All You Need
- Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain
- Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanism
- Situat3DChange: Situated 3D Change Understanding Dataset for Multimodal Large Language Model
- Part II: ROLL Flash -- Accelerating RLVR and Agentic Training with Asynchrony
- FOSSIL: Harnessing Feedback on Suboptimal Samples for Data-Efficient Generalisation with Imitation Learning for Embodied Vision-and-Language Tasks
- PADME: Procedure Aware DynaMic Execution
- ESCA: Contextualizing Embodied Agents via Scene-Graph Generation
- SkillOS: Learning Skill Curation for Self-Evolving Agents
- Can RL Improve Generalization of LLM Agents? An Empirical Study
- Dyna-Mind: Learning to Simulate from Experience for Better AI Agents
- Agent Learning via Early Experience
- Learning on the Job: An Experience-Driven Self-Evolving Agent for Long-Horizon Tasks
- MIMIC: Integrating Diverse Personality Traits for Better Game Testing Using Large Language Model
- Information Seeking for Robust Decision Making under Partial Observability
- Adaptive Tool Generation with Models as Tools and Reinforcement Learning
- The Cognitive Bandwidth Bottleneck: Shifting Long-Horizon Agent from Planning with Actions to Planning with Schemas
- Constrained Natural Language Action Planning for Resilient Embodied Systems
- LexiCon: a Benchmark for Planning under Temporal Constraints in Natural Language
- A Goal Without a Plan Is Just a Wish: Efficient and Effective Global Planner Training for Long-Horizon Agent Tasks
- AgentRL: Scaling Agentic Reinforcement Learning with a Multi-Turn, Multi-Task Framework
- Learning Efficient Guardrails for Compliance
- Fine-tuning with RAG for Improving LLM Learning of New Skills
- A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning
- Memory-Driven Self-Improvement for Decision Making with Large Language Models
- Where LLM Agents Fail and How They can Learn From Failures
- MemGen: Weaving Generative Latent Memory for Self-Evolving Agents
- AutoContext: Instance-Level Context Learning for LLM Agents
- ByteSized32Refactored: Towards an Extensible Interactive Text Games Corpus for LLM World Modeling and Evaluation
- Rethinking Reward Miscalibration of GRPO in Agentic RL
- Measuring Physical-World Privacy Awareness of Large Language Models: An Evaluation Benchmark
- LAGEA: Language Guided Embodied Agents for Robotic Manipulation
- Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning
- EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning
- Solving the Granularity Mismatch: Hierarchical Preference Learning for Long-Horizon LLM Agents
- Vision Language Models Cannot Plan, but Can They Formalize?
- Training Task Reasoning LLM Agents for Multi-turn Task Planning via Single-turn Reinforcement Learning
- From Scoring to Acting: Outcome-Verified Comparative Self-Distillation for LLM Agents
- LabEvolver: Training-Free Experience Evolution for Safe and Grounded Wet-Lab Agents
- Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation
- TAPO: Transition-Aware Policy Optimization for LLM Agents
- Group-Reflective Self-Distillation for Agentic Reinforcement Learning
- SKILL-KD: Contrastive Skill Distillation for LLM Agents
- Leveraging Trajectory Graphs for Pre-Execution Error Diagnosis in Agentic LLM Systems
- Bridging Inference-Time Scaling and Episodic Memory with Action-Centric Graphs
- DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes
- The Artificial Intelligence Cognitive Examination: A Survey on the Evolution of Multimodal Evaluation From Recognition to Reasoning
- Reflect before Act: Proactive Error Correction in Language Models
- Code Driven Planning with Domain-Adaptive Critic
- Generalizable End-to-End Tool-Use RL with Synthetic CodeGym
- MCTS-EP: Empowering Embodied Planning with Online Preference Optimization
- Modeling Worlds in Text
- Generalizability of Large Language Model-Based Agents: A Comprehensive Survey
- From Language to Action: A Review of Large Language Models as Autonomous Agents and Tool Users
- H2R: Hierarchical Hindsight Reflection for Multi-Task LLM Agents
- Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents
- How well can LLMs provide planning feedback in grounded environments?
- AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
- Meta-Policy Reflexion: Reusable Reflective Memory and Rule Admissibility for Resource-Efficient LLM Agent
- World Model Implanting for Test-time Adaptation of Embodied Agents
- Succeed or Learn Slowly: Sample Efficient Off-Policy Reinforcement Learning for Mobile App Control
- HiPlan: Hierarchical Planning for LLM-Based Agents with Adaptive Global-Local Guidance
- Language and Experience: A Computational Model of Social Learning in Complex Tasks
- Coarse-to-Fine Grounded Memory for LLM Agent Planning
- Prompt Orchestration Markup Language
- HeroBench: A Benchmark for Long-Horizon Planning and Structured Reasoning in Virtual Worlds
- Leveraging OS-Level Primitives for Robotic Action Management
- OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows
- Intrinsic Memory Agents: Heterogeneous Multi-Agent LLM Systems through Structured Contextual Memory
- GVGAI-LLM: Evaluating Large Language Model Agents with Infinite Games
- Memp: Exploring Agent Procedural Memory
- OmniEAR: Benchmarking Agent Reasoning in Embodied Tasks
- OmniPlay: Benchmarking Omni-Modal Models on Omni-Modal Game Playing
- Enhancing Vision-Language Model Training with Reinforcement Learning in Synthetic Worlds for Real-World Success
- RCR-Router: Efficient Role-Aware Context Routing for Multi-Agent LLM Systems with Structured Memory
- CookBench: A Long-Horizon Embodied Planning Benchmark for Complex Cooking Scenarios
- Adaptive Command: Real-Time Policy Adjustment via Language Models in StarCraft II
- Beyond Policy Optimization: A Data Curation Flywheel for Sparse-Reward Long-Horizon Planning
- L3M+P: Lifelong Planning with Large Language Models
- Sari Sandbox: A Virtual Retail Store Environment for Embodied AI Agents
- PilotRL: Training Language Model Agents via Global Planning-Guided Progressive Reinforcement Learning
- Blueprint First, Model Second: A Framework for Deterministic LLM Workflow
- DICE: Dynamic In-Context Example Selection in LLM Agents via Efficient Knowledge Transfer
- CoEx -- Co-evolving World-model and Exploration
- MIRAGE-Bench: LLM Agent is Hallucinating and Where to Find Them
- MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models
- Agentic Reinforced Policy Optimization
- Weak-to-Strong Generalization with Failure Trajectories: A Tree-based Approach to Elicit Optimal Policy in Strong Models
- FCRF: Flexible Constructivism Reflection for Long-Horizon Robotic Task Planning with Large Language Models
- AgentFly: Extensible and Scalable Reinforcement Learning for LM Agents
- PivotRL: High Accuracy Agentic Post-Training at Low Compute Cost
- Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments
- A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning
- Graph World Model
- Towards Agentic RAG with Deep Reasoning: A Survey of RAG-Reasoning Systems in LLMs
- SAND: Boosting LLM Agents with Self-Taught Action Deliberation
- CRISP: Complex Reasoning with Interpretable Step-based Plans
- NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial Observability
- ECom-Bench: Can LLM Agent Resolve Real-World E-commerce Customer Support Issues?
- DASH-OPD: Discrepancy-Aware Switching with Hysteresis for On-Policy Distillation
- Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution
- Self-Supervised Skill Optimization
- APPO: Agentic Procedural Policy Optimization
- Learning Stateful Predictive Knowledge From Experience
- Contextual Experience Replay for Self-Improvement of Language Agents
- Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning
- Universal Retrieval for Multimodal Trajectory Modeling
- SEEA-R1: Tree-Structured Reinforcement Fine-Tuning for Self-Evolving Embodied Agents
- World-aware Planning Narratives Enhance Large Vision-Language Model Planner
- G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems
- OmniReflect: Discovering Transferable Constitutions for LLM agents via Neuro-Symbolic Reflections
- SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents
- DualTHOR: A Dual-Arm Humanoid Simulation Platform for Contingency-Aware Planning
- From Passive to Active Reasoning: Can Large Language Models Ask the Right Questions under Incomplete Information?
- Unveiling the Learning Mind of Language Models: A Cognitive Framework and Empirical Study
- Towards Pervasive Distributed Agentic Generative AI -- A State of The Art
- Leveraging In-Context Learning for Language Model Agents
- Hierarchical Task Learning from Language Instructions with Unified Transformers and Self-Monitoring
- Wide-Horizon Thinking and Simulation-Based Evaluation for Real-World LLM Planning with Multifaceted Constraints
- IndoorWorld: Integrating Physical Task Solving and Social Simulation in A Heterogeneous Multi-Agent Environment
- Efficient LLM Collaboration via Planning
- OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems
- Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives
- AgentSwift: Efficient LLM Agent Design via Value-guided Hierarchical Search
- From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems
- SocialAI: Benchmarking Socio-Cognitive Abilities in Deep Reinforcement Learning Agents
- Enhancing Decision-Making of Large Language Models via Actor-Critic
- PGPO: Enhancing Agent Reasoning via Pseudocode-style Planning Guided Preference Optimization
- Divide, Optimize, Merge: Fine-Grained LLM Agent Optimization at Scale
- ARIA: Training Language Agents with Intention-Driven Reward Aggregation
- Deep Research Bench: Evaluating AI Web Research Agents
- From Knowledge to Noise: CTIM-Rover and the Pitfalls of Episodic Memory in Software Engineering Agents
- Topological Structure Learning Should Be A Research Priority for LLM-Based Multi-Agent Systems
- 3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Model
- SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
- Make Planning Research Rigorous Again!
- Large Language Models for Planning: A Comprehensive and Systematic Survey
- Multi-Agent Collaboration via Evolving Orchestration
- Training LLM-Based Agents with Synthetic Self-Reflected Trajectories and Partial Masking
- EMAC+: Embodied Multimodal Agent for Collaborative Planning with VLM+LLM
- Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories
- VideoGameBench: Can Vision-Language Models complete popular video games?
- PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning
- FedWorld: Scope-Aware Federation of Agent World Models
- The Real Barrier to LLM Agent Usability is Agentic ROI
- Runaway is Ashamed, But Helpful: On the Early-Exit Behavior of Large Language Model-based Agents in Embodied Environments
- Distilling LLM Agent into Small Models with Retrieval and Code Tools
- Trinity-RFT: A General-Purpose and Unified Framework for Reinforcement Fine-Tuning of Large Language Models
- SkillTrace: Traversing a Query-Skill Graph for Composable LLM Agents
- T1: A Tool-Oriented Conversational Dataset for Multi-Turn Agentic Planning
- Look Ahead Before You Distill: Future Trajectory Validation of Teacher Guidance for Agentic On-Policy Distillation
- MCP-RADAR: A Multi-Dimensional Benchmark for Evaluating Tool Use Capabilities in Large Language Models
- From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization
- MemArbiter: Decision-Time Memory Arbitration for Long-Horizon LLM Agents
- Beyond Needle(s) in the Embodied Haystack: Environment, Architecture, and Training Considerations for Long Context Reasoning
- ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State Reflection
- Structured Agent Distillation for Large Language Model
- Cost-Awareness in Tree-Search LLM Planning: A Systematic Study
- Zero-Shot Iterative Formalization and Planning in Partially Observable Environments
- Long-Horizon Embodied Decision-Making via Multimodal Memory Compression
- ALAS: A Stateful Multi-LLM Agent Framework for Disruption-Aware Planning
- Retrospex: Language Agent Meets Offline Reinforcement Learning Critic
- LLM-BABYBENCH: Understanding and Evaluating Grounded Planning and Reasoning in LLMs
- ProxyPrompt: Securing System Prompts against Prompt Extraction Attacks
- AgentSLABench: Evaluating and Benchmarking Agentic Systems Under Resource Constraints
- AdvPlan-Bench: Adversarial Evaluation of Structured Plan-Generation Agents
- Group-in-Group Policy Optimization for LLM Agent Training
- CrystalMem: Elastic Memory for Self-Evolving LLM Agents via Knowledge Crystallization
- Cache-Efficient Posterior Sampling for Reinforcement Learning with LLM-Derived Priors Across Discrete and Continuous Domains
- Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
- Can LLM Agents Be CFOs? Benchmarking Long-Horizon Resource Allocation in an Uncertain Enterprise Environment
- Exploration and Exploitation Errors Are Measurable for Language Model Agents
- Graph-based Agent Memory: Taxonomy, Techniques, and Applications
- Rethinking Self-Evolving Agent Skills: Feedback Dynamics over Multiple Rounds
- Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking
- SIMMER: Benchmarking Latent Failures in LLM Executable Planning with a World Model
- The Synthetic Web: Adversarially-Curated Mini-Internets for Diagnosing Epistemic Weaknesses of Language Agents
- Towards Efficient Online Tuning of VLM Agents via Counterfactual Soft Reinforcement Learning
- Self-Generated In-Context Examples Improve LLM Agents for Sequential Decision-Making Tasks
- Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning
- AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution
- SkillNet: Create, Evaluate, and Connect AI Skills
- SkillsInjector: Dynamic Skill Context Construction for LLM Agents
- PatchBoard: Schema-Grounded State Mutation for Reliable and Auditable LLM Multi-Agent Collaboration
- Hierarchical Prompt-Domain Control and Learning for Resource-Constrained Agentic Language Models
- Self-Distilled Agentic Reinforcement Learning
- Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM Agents
- PRISM: Perception Reasoning Interleaved for Sequential Decision Making
- Large Language Model Agents Are Not Always Faithful Self-Evolvers
- Evolution of Cooperation in LLM-Agent Societies: A Preliminary Study Using Different Punishment Strategies
- Generative AI in Embodied Systems: System-Level Analysis of Performance, Efficiency and Scalability
- Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks
- Auto-SLURP: A Benchmark Dataset for Evaluating Multi-Agent Frameworks in Smart Personal Assistant
- The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break
- SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
- Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills
- Verifiable Memory: Learning Unified Memory Management with Local and Global Verifiers for Large Language Model Agents
- Agentic Reinforcement Learning with Self-Distilled Reward Shaping
- MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory
- Improving Large Language Model Planning with Action Sequence Similarity
- InsightEmb: Learning Action-Intent Embeddings for Agentic Insight Retrieval
- Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation
- State2State: Environment-Derived Mid-Training for LLM Agents
- Strategic Evaluation of Planning Strategies for LLM Agents in Cyber-Physical Systems
- Tree of Thoughts as a Classical Heuristic Search Problem: Formal Foundations and Design Patterns
- Monte Carlo Planning with Large Language Model for Text-Based Game Agents
- WALL-E 2.0: World Alignment by NeuroSymbolic Learning improves World Model-based LLM Agents
- PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities
- InstructRAG: Leveraging Retrieval-Augmented Generation on Instruction Graphs for LLM-Based Task Planning
- SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries
- When Do Prompt-Side Agent Playbooks Transfer? Accuracy, Cost, and Runtime Shift in Agent Deployment
- AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
- Seeing Is Not Deciding: Can Multimodal LLMs Act as Effective CEOs?
- EvoHarness-RL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents
- When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents
- Breaking the Data Barrier -- Building GUI Agents Through Task Generalization
- A Desideratum for Conversational Agents: Capabilities, Challenges, and Future Directions
Related