A Generalist Agent
2022/05/12 by Scott Reed, Konrad Zolna, Emilio Parisotto +17 · 1 voice · 150 citations
Computer Science · #cs.AI #cs.CL #cs.LG #cs.RO
paper · pdf
published as Transactions on Machine Learning Research, 11/2022, https://openreview.net/forum?id=1ikK0kHjvj · Published at TMLR, 42 pages
arxiv created 2022/11/11 · arxiv updated 2022/11/14
Abstract
Inspired by progress in large-scale language modeling, we apply a similar approach towards building a single generalist agent beyond the realm of text outputs. The agent, which we refer to as Gato, works as a multi-modal, multi-task, multi-embodiment generalist policy. The same network with the same weights can play Atari, caption images, chat, stack blocks with a real robot arm and much more, deciding based on its context whether to output text, joint torques, button presses, or other tokens. In this report we describe the model and the data, and document the current capabilities of Gato.
Cited by
- SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation
- Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation
- Patch Policy: Efficient Embodied Control via Dense Visual Representations
- Do Maps Still Matter for Machines: Revisiting the Role of Choropleth Maps in Foundation Model Spatial Understanding
- TempoVLA: Learning Speed-Controllable Vision-Language-Action Policies
- PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution
- SportD: Can VLMs Physically Strategize?
- Active Real-World Factor-Based Evaluation for Generalist Robot Policies
- Generalist AI control: Towards multi-purpose adaptive algorithms
- Why AI systems don't learn and what to do about it: Lessons on autonomous learning from cognitive science
- Arnold: A multi-task, multi-embodiment muscle transformer policy
- What does really matter in image goal navigation?
- Human-Curated Data Authoring with LLMs: A Small-Data Approach to Domain Adaptation
- General agents contain world models
- RANa: Retrieval-Augmented Navigation
- VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models
- Role-Based Fault Tolerance System for LLM RL Post-Training
- The Curse of Precision: A Data Scaling Law for High-Precision Robotic Manipulation
- Bridging MARL to SARL: An Order-Independent Multi-Agent Transformer via Latent Consensus
- The Cartesian Cut in Agentic AI
- mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
- OXE-AugE: A Large-Scale Robot Augmentation of OXE for Scaling Cross-Embodiment Policy Learning
- An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
- Invariance Co-training for Robot Visual Generalization
- SIMA 2: A Generalist Embodied Agent for Virtual Worlds
- FALCON: Actively Decoupled Visuomotor Policies for Loco-Manipulation with Foundation-Model-Based Coordination
- The brain-AI convergence: Predictive and generative world models for general-purpose computation
- Cross-Geometry Transfer Learning in Fast Electromagnetic Shower Simulation
- Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
- Learning Massively Multitask World Models for Continuous Control
- Compressor-VLA: Instruction-Guided Visual Token Compression for Efficient Robotic Manipulation
- EchoVLA: Synergistic Declarative Memory for VLA-Driven Mobile Manipulation
- IPR-1: Interactive Physical Reasoner
- AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models
- Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
- Representation Learning Enables Scalable Multitask Deep Reinforcement Learning
- Balance Equation-based Distributionally Robust Offline Imitation Learning
- Balancing Multi-modal Sensor Learning via Multi-objective Optimization
- How Do VLAs Effectively Inherit from VLMs?
- Dexterous Robotic Piano Playing at Scale
- How Far Are Surgeons from Surgical World Models? A Pilot Study on Zero-shot Surgical Video Generation with Expert Assessment
- UMI-on-Air: Embodiment-Aware Guidance for Embodiment-Agnostic Visuomotor Policies
- Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning
- Integrative neurocybernetic modeling in the era of large-scale neuroscience
- PFEA: An LLM-based High-Level Natural Language Planning and Feedback Embodied Agent for Human-Centered AI
- Human Machine Social Hybrid Intelligence:A Collaborative Decision Making Framework for Large Model Agent Groups and Human Experts
- Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents
- From Masks to Worlds: A Hitchhiker's Guide to World Models
- Socialized Learning and Emergent Behaviors in Multi-Agent Systems based on Multimodal Large Language Models
- Consistent Zero-Shot Imitation with Contrastive Goal Inference
- Plasma Shape Control via Zero-shot Generative Reinforcement Learning
- End-to-end Listen, Look, Speak and Act
- Combining Reinforcement Learning and Behavior Trees for NPCs in Video Games with AMD Schola
- Towards Neurocognitive-Inspired Intelligence: From AI's Structural Mimicry to Human-Like Functional Cognition
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- EmbodiedCoder: Parameterized Embodied Mobile Manipulation via Modern Coding Model
- D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI
- Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI
- Seeing Space and Motion: Enhancing Latent Actions with Spatial and Dynamic Awareness for VLA
- Accelerating Transformers in Online RL
- Indirect Attention: Turning Context Misalignment into a Feature
- Fidelity-Aware Data Composition for Robust Robot Generalization
- IA-VLA: Input Augmentation for Vision-Language-Action models in settings with semantically complex tasks
- In-Context Compositional Q-Learning for Offline Reinforcement Learning
- LocoFormer: Generalist Locomotion via Long-context Adaptation
- Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models
- Pixel Motion Diffusion is What We Need for Robot Control
- Embodied AI: From LLMs to World Models
- A scalar per patch from pre-trained ViTs enables fast moving navigation in the real world
- Algebraic Approach to Ridge-Regularized Mean Squared Error Minimization in Minimal ReLU Neural Network
- OmniVLA: An Omni-Modal Vision-Language-Action Model for Robot Navigation
- Pure Vision Language Action (VLA) Models: A Comprehensive Survey
- Self-Improving Embodied Foundation Models
- GeoAware-VLA: Implicit Geometry Aware Vision-Language-Action Model
- Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges
- A Generalisable Generative Model for Multi-Detector Calorimeter Simulation
- A Data-Driven Discretized CS:GO Simulation Environment to Facilitate Strategic Multi-Agent Planning Research
- ProfilingAgent: Profiling-Guided Agentic Reasoning for Adaptive Model Optimization
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
- Data Retrieval with Importance Weights for Few-Shot Imitation Learning
- VistaWise: Building Cost-Effective Agent with Cross-Modal Knowledge Graph for Minecraft
- Survey of Vision-Language-Action Models for Embodied Manipulation
- Human Centric General Physical Intelligence for Agile Manufacturing Automation
- Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges
- The Othello AI Arena: Evaluating Intelligent Systems Through Limited-Time Adaptation to Unseen Boards
- Vision Generalist Model: A Survey
- DeepFleet: Multi-Agent Foundation Models for Mobile Robots
- Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation
- In-Context Reinforcement Learning via Communicative World Models
- ASkDAgger: Active Skill-level Data Aggregation for Interactive Imitation Learning
- Efficient Morphology-Aware Policy Transfer to New Embodiments
- Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies
- RICL: Adding In-Context Adaptability to Pre-Trained Vision-Language-Action Models
- Frequency Point Game Environment for UAVs via Expert Knowledge and Large Language Model
- COLLAGE: Adaptive Fusion-based Retrieval for Augmented Policy Learning
- UniLegs: Universal Multi-Legged Robot Control through Morphology-Agnostic Policy Distillation
- Exploring the Link Between Bayesian Inference and Embodied Intelligence: Toward Open Physical-World Embodied AI Systems
- GR-3 Technical Report
- Your Attention Matters: to Improve Model Robustness to Noise and Spurious Correlations
- Automatic Generation of High-Performance RL Environments
- Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities
- Hybrid Reasoning for Perception, Explanation, and Autonomous Action in Manufacturing
- Foundation Model Driven Robotics: A Comprehensive Review
- Value from Observations: Towards Large-Scale Imitation Learning via Self-Improvement
- A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation
- DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
- Grounding Intelligence in Movement
- ACTLLM: Action Consistency Tuned Large Language Model
- Generalizable Agent Modeling for Agent Collaboration-Competition Adaptation with Multi-Retrieval and Dynamic Generation
- FlowRAM: Grounding Flow Matching Policy with Region-Aware Mamba Framework for Robotic Manipulation
- Play to Generalize: Learning to Reason Through Game Play
- TrojanTO: Action-Level Backdoor Attacks against Trajectory Optimization Models
- Perspective on Utilizing Foundation Models for Laboratory Automation in Materials Research
- Efficient Multi-Camera Tokenization with Triplanes for End-to-End Driving
- Robust In-Context Reinforcement Learning Under Reward Poisoning Attacks
- Demonstrating Multi-Suction Item Picking at Scale via Multi-Modal Learning of Pick Success
- VITA: Zero-Shot Value Functions via Test-Time Adaptation of Vision-Language Models
- Dynamic Mixture of Progressive Parameter-Efficient Expert Library for Lifelong Robot Learning
- Benchmarking Vision, Language, & Action Models in Procedurally Generated, Open Ended Action Environments
- UniCO: Towards a Unified Model for Combinatorial Optimization Problems
- AI Agent Behavioral Science
- Horizon Reduction Makes RL Scalable
- Mastering Massive Multi-Task Reinforcement Learning via Mixture-of-Expert Decision Transformer
- Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners
- ReFineVLA: Reasoning-Aware Teacher-Guided Transfer Fine-Tuning
- Filtering Learning Histories Enhances In-Context Reinforcement Learning
- AnyBody: A Benchmark Suite for Cross-Embodiment Manipulation
- Structured Agent Distillation for Large Language Model
- APEX: Empowering LLMs with Physics-Based Task Planning for Real-time Insight
- Building spatial world models from sparse transitional episodic memories
- Tool-Aided Evolutionary LLM for Generative Policy Toward Efficient Resource Management in Wireless Federated Learning
- AI Agents vs. Agentic AI: A Conceptual Taxonomy, Applications and Challenges
- Hierarchical Surgical Robot Transformer (SRT-H): Imitation Learning for Autonomous Surgery
- RAI: Flexible Agent Framework for Embodied AI
- Pixel Motion as Universal Representation for Robot Control
- Factorization Regret mediates compositional generalization in latent space
- PainFormer: a Vision Foundation Model for Automatic Pain Assessment
- Pretrained Vision-Language-Action Models are Surprisingly Resistant to Forgetting in Continual Learning
- A Survey of Interactive Generative Video
- Act-Observe-Rewrite: Multimodal Coding Agents as In-Context Policy Learners for Robot Manipulation
- From Kepler to Newton: Inductive Biases Guide Learned World Models in Transformers
- RL-Driven Data Generation for Robust Vision-Based Dexterous Grasping
- LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities
- Human-Level Competitive Pokémon via Scalable Offline Reinforcement Learning with Transformers
- VLMs for Videogame Data Annotation
- Robust Scene Transfer for PointGoal Navigation via Privileged Sensor Guided Contrastive Learning
- State Estimation Using Particle Filtering in Adaptive Machine Learning Methods: Integrating Q-Learning and NEAT Algorithms with Noisy Radar Measurements
- Bitter lesson [wikipedia]
- Gato (DeepMind) [wikipedia]
- Generative AI [wikipedia]
Discussions
Related