A Generalist Agent
2022/05/12 by Scott Reed, Konrad Zolna, Emilio Parisotto +17 · 1 voice · 100 citations
#cs.AI #cs.CL #cs.LG #cs.RO
paper · pdf
Abstract
Inspired by progress in large-scale language modeling, we apply a similar approach towards building a single generalist agent beyond the realm of text outputs. The agent, which we refer to as Gato, works as a multi-modal, multi-task, multi-embodiment generalist policy. The same network with the same weights can play Atari, caption images, chat, stack blocks with a real robot arm and much more, deciding based on its context whether to output text, joint torques, button presses, or other tokens. In this report we describe the model and the data, and document the current capabilities of Gato.
Cited by
- SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation
- Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation
- Patch Policy: Efficient Embodied Control via Dense Visual Representations
- Do Maps Still Matter for Machines: Revisiting the Role of Choropleth Maps in Foundation Model Spatial Understanding
- TempoVLA: Learning Speed-Controllable Vision-Language-Action Policies
- PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution
- SportD: Can VLMs Physically Strategize?
- Active Real-World Factor-Based Evaluation for Generalist Robot Policies
- Generalist AI control: Towards multi-purpose adaptive algorithms
- Why AI systems don't learn and what to do about it: Lessons on autonomous learning from cognitive science
- Arnold: a generalist muscle transformer policy
- What does really matter in image goal navigation?
- Simulating Society Requires Simulating Thought
- General agents contain world models
- RANa: Retrieval-Augmented Navigation
- VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models
- Role-Based Fault Tolerance System for LLM RL Post-Training
- The Curse of Precision: A Data Scaling Law for High-Precision Robotic Manipulation
- Bridging MARL to SARL: An Order-Independent Multi-Agent Transformer via Latent Consensus
- The Cartesian Cut in Agentic AI
- mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
- OXE-AugE: A Large-Scale Robot Augmentation of OXE for Scaling Cross-Embodiment Policy Learning
- An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
- Invariance Co-training for Robot Visual Generalization
- SIMA 2: A Generalist Embodied Agent for Virtual Worlds
- FALCON: Actively Decoupled Visuomotor Policies for Loco-Manipulation with Foundation-Model-Based Coordination
- The brain-AI convergence: Predictive and generative world models for general-purpose computation
- Cross-Geometry Transfer Learning in Fast Electromagnetic Shower Simulation
- Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
- Learning Massively Multitask World Models for Continuous Control
- Compressor-VLA: Instruction-Guided Visual Token Compression for Efficient Robotic Manipulation
- EchoVLA: Synergistic Declarative Memory for VLA-Driven Mobile Manipulation
- IPR-1: Interactive Physical Reasoner
- AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models
- Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
- Representation Learning Enables Scalable Multitask Deep Reinforcement Learning
- Balance Equation-based Distributionally Robust Offline Imitation Learning
- Balancing Multi-modal Sensor Learning via Multi-objective Optimization
- How Do VLAs Effectively Inherit from VLMs?
- Dexterous Robotic Piano Playing at Scale
- How Far Are Surgeons from Surgical World Models? A Pilot Study on Zero-shot Surgical Video Generation with Expert Assessment
- UMI-on-Air: Embodiment-Aware Guidance for Embodiment-Agnostic Visuomotor Policies
- Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning
- Integrative neurocybernetic modeling in the era of large-scale neuroscience
- PFEA: An LLM-based High-Level Natural Language Planning and Feedback Embodied Agent for Human-Centered AI
- Human Machine Social Hybrid Intelligence:A Collaborative Decision Making Framework for Large Model Agent Groups and Human Experts
- Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents
- From Masks to Worlds: A Hitchhiker's Guide to World Models
- Socialized Learning and Emergent Behaviors in Multi-Agent Systems based on Multimodal Large Language Models
- Consistent Zero-Shot Imitation with Contrastive Goal Inference
- Plasma Shape Control via Zero-shot Generative Reinforcement Learning
- End-to-end Listen, Look, Speak and Act
- Combining Reinforcement Learning and Behavior Trees for NPCs in Video Games with AMD Schola
- Towards Neurocognitive-Inspired Intelligence: From AI's Structural Mimicry to Human-Like Functional Cognition
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- EmbodiedCoder: Parameterized Embodied Mobile Manipulation via Modern Coding Model
- D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI
- Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI
- Seeing Space and Motion: Enhancing Latent Actions with Spatial and Dynamic Awareness for VLA
- Accelerating Transformers in Online RL
- Indirect Attention: Turning Context Misalignment into a Feature
- Fidelity-Aware Data Composition for Robust Robot Generalization
- IA-VLA: Input Augmentation for Vision-Language-Action models in settings with semantically complex tasks
- In-Context Compositional Q-Learning for Offline Reinforcement Learning
- LocoFormer: Generalist Locomotion via Long-context Adaptation
- Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models
- Pixel Motion Diffusion is What We Need for Robot Control
- Embodied AI: From LLMs to World Models
- A scalar per patch from pre-trained ViTs enables fast moving navigation in the real world
- Algebraic Approach to Ridge-Regularized Mean Squared Error Minimization in Minimal ReLU Neural Network
- OmniVLA: An Omni-Modal Vision-Language-Action Model for Robot Navigation
- Pure Vision Language Action (VLA) Models: A Comprehensive Survey
- Self-Improving Embodied Foundation Models
- GeoAware-VLA: Implicit Geometry Aware Vision-Language-Action Model
- Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges
- A Generalisable Generative Model for Multi-Detector Calorimeter Simulation
- A Data-Driven Discretized CS:GO Simulation Environment to Facilitate Strategic Multi-Agent Planning Research
- ProfilingAgent: Profiling-Guided Agentic Reasoning for Adaptive Model Optimization
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
- Data Retrieval with Importance Weights for Few-Shot Imitation Learning
- VistaWise: Building Cost-Effective Agent with Cross-Modal Knowledge Graph for Minecraft
- Survey of Vision-Language-Action Models for Embodied Manipulation
- Human Centric General Physical Intelligence for Agile Manufacturing Automation
- Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges
- The Othello AI Arena: Evaluating Intelligent Systems Through Limited-Time Adaptation to Unseen Boards
- DeepFleet: Multi-Agent Foundation Models for Mobile Robots
- Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation
- In-Context Reinforcement Learning via Communicative World Models
- ASkDAgger: Active Skill-level Data Aggregation for Interactive Imitation Learning
- Efficient Morphology-Aware Policy Transfer to New Embodiments
- Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies
- RICL: Adding In-Context Adaptability to Pre-Trained Vision-Language-Action Models
- Frequency Point Game Environment for UAVs via Expert Knowledge and Large Language Model
- COLLAGE: Adaptive Fusion-based Retrieval for Augmented Policy Learning
- UniLegs: Universal Multi-Legged Robot Control through Morphology-Agnostic Policy Distillation
- Exploring the Link Between Bayesian Inference and Embodied Intelligence: Toward Open Physical-World Embodied AI Systems
- Your Attention Matters: to Improve Model Robustness to Noise and Spurious Correlations
- Bitter lesson [wikipedia]
- Gato (DeepMind) [wikipedia]
- Generative AI [wikipedia]
Discussions
Related