ProcTHOR: Large-Scale Embodied AI Using Procedural Generation
2022/06/14 by Matt Deitke, Deitke, Matt, Eli VanderBilt +19 · 1 voice · 126 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Robotics (cs.RO) #cs.AI #cs.CV #cs.RO
paper · pdf · doi:10.48550/arxiv.2206.06994
ProcTHOR website: https://procthor.allenai.org
arxiv created 2022/06/14 · arxiv published 2022/06/14 · arxiv updated 2022/06/15
Abstract
Massive datasets and high-capacity models have driven many recent advancements in computer vision and natural language understanding. This work presents a platform to enable similar success stories in Embodied AI. We propose ProcTHOR, a framework for procedural generation of Embodied AI environments. ProcTHOR enables us to sample arbitrarily large datasets of diverse, interactive, customizable, and performant virtual environments to train and evaluate embodied agents across navigation, interaction, and manipulation tasks. We demonstrate the power and potential of ProcTHOR via a sample of 10,000 generated houses and a simple neural model. Models trained using only RGB images on ProcTHOR, with no explicit mapping and no human task supervision produce state-of-the-art results across 6 embodied AI benchmarks for navigation, rearrangement, and arm manipulation, including the presently running Habitat 2022, AI2-THOR Rearrangement 2022, and RoboTHOR challenges. We also demonstrate strong 0-shot results on these benchmarks, via pre-training on ProcTHOR with no fine-tuning on the downstream benchmark, often beating previous state-of-the-art systems that access the downstream training data.
Cited by
- VULCAN: Tool-Augmented Multi Agents for Iterative 3D Object Arrangement
- PlanCraft: Sketch, Refine, and Furnish for Architect-Inspired Progressive 3D Residential Scene Generation
- DSCD-Nav: Dual-Stance Cooperative Debate for Object Navigation
- TongSIM: A General Platform for Simulating Intelligent Machines
- DeliveryBench: Can Agents Earn Profit in Real World?
- M3-Verse: A "Spot the Difference" Challenge for Large Multimodal Models
- Scaling Spatial Reasoning in MLLMs through Programmatic Data Synthesis
- Sceniris: A Fast Procedural Scene Generation Framework
- Robust Single-shot Structured Light 3D Imaging via Neural Feature Decoding
- Toward Ambulatory Vision: Learning Visually-Grounded Active View Selection
- D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and Navigation
- RoomPilot: Controllable Indoor Scene Synthesis via Multimodal Semantic Parsing
- UrbanNav: Learning Language-Guided Urban Navigation from Web-Scale Human Trajectories
- SIMA 2: A Generalist Embodied Agent for Virtual Worlds
- MarketGen: A Scalable Simulation Platform with Auto-Generated Embodied Supermarket Environments
- BRIC: Bridging Kinematic Plans and Physical Control at Test Time
- Thinking in 360°: Humanoid Visual Search in the Wild
- Disc3D: Automatic Curation of High-Quality 3D Dialog Data via Discriminative Object Referring
- InfiniBench: Infinite Benchmarking for Visual Spatial Reasoning with Customizable Scene Complexity
- EchoVLA: Synergistic Declarative Memory for VLA-Driven Mobile Manipulation
- TP-MDDN: Task-Preferenced Multi-Demand-Driven Navigation with Autonomous Decision-Making
- POMA-3D: The Point Map Way to 3D Scene Understanding
- YOWO: You Only Walk Once to Jointly Map An Indoor Scene and Register Ceiling-mounted Cameras
- Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
- Cambrian-S: Towards Spatial Supersensing in Video
- SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding
- LEGO-Eval: Towards Fine-Grained Evaluation on Synthesizing 3D Embodied Environments with Tool Augmentation
- StructureGS: Structure-aware Gaussian Splatting for Articulated Object Reconstruction
- BioVLN: A Simulation Platform for Visual Language Navigation in Biomedical Laboratories
- PhyScensis: Physics-Augmented LLM Agents for Complex Physical Scene Arrangement
- Demeter: A Parametric Model of Crop Plant Morphology from the Real World
- GaussGym: An open-source real-to-sim framework for learning locomotion from pixels
- UrbanVerse: Scaling Urban Simulation by Watching City-Tour Videos
- IL3D: A Large-Scale Indoor Layout Dataset for LLM-Driven 3D Scene Generation
- EmboMatrix: A Scalable Training-Ground for Embodied Decision-Making
- CompassNav: Steering From Path Imitation To Decision Understanding In Navigation
- Towards Proprioception-Aware Embodied Planning for Dual-Arm Humanoid Robots
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- MetaFind: Scene-Aware 3D Asset Retrieval for Coherent Metaverse Scene Generation
- Text-to-Scene with Large Reasoning Models
- M3DLayout: A Multi-Source Dataset of 3D Indoor Layouts and Structured Descriptions for 3D Generation
- Context and Diversity Matter: The Emergence of In-Context Learning in World Models
- SemSight: Probabilistic Bird's-Eye-View Prediction of Multi-Level Scene Semantics for Navigation
- PersONAL: Towards a Comprehensive Benchmark for Personalized Embodied Agents
- SceneWeaver: All-in-One 3D Scene Synthesis with an Extensible and Self-Reflective Agent
- Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation
- HLG: Comprehensive 3D Room Construction via Hierarchical Layout Generation
- Generalizable Domain Adaptation for Sim-and-Real Policy Co-Training
- Investigating Domain Gaps for Indoor 3D Object Detection
- Embodied Navigation Foundation Model
- GBPP: Grasp-Aware Base Placement Prediction for Robots via Two-Stage Learning
- InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts
- HumanoidVerse: A Versatile Humanoid for Vision-Language Guided Multi-Object Rearrangement
- TopoNav: Topological Graphs as a Key Enabler for Advanced Object Navigation
- SemLayoutDiff: Semantic Layout Generation with Diffusion Model for Indoor Scene Synthesis
- Virtual Community: An Open World for Humans, Robots, and Society
- ResPlan: A Large-Scale Vector-Graph Dataset of 17,000 Residential Floor Plans
- OVSegDT: Segmenting Transformer for Open-Vocabulary Object Goal Navigation
- AgentWorld: An Interactive Simulation Platform for Scene Construction and Mobile Robotic Manipulation
- Learning Robust Intervention Representations with Delta Embeddings
- Rethinking the Embodied Gap in Vision-and-Language Navigation: A Holistic Study of Physical and Visual Disparities
- CookBench: A Long-Horizon Embodied Planning Benchmark for Complex Cooking Scenarios
- L3M+P: Lifelong Planning with Large Language Models
- Interleaved LLM and Motion Planning for Generalized Multi-Object Collection in Large Scene Graphs
- HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels
- OctoNav: Towards Generalist Embodied Navigation
- From Scan to Action: Leveraging Realistic Scans for Embodied Scene Understanding
- CogDDN: A Cognitive Demand-Driven Navigation with Decision Optimization and Dual-Process Thinking
- 3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds
- LiteReality: Graphics-Ready 3D Scene Reconstruction from RGB-D Scans
- Evaluating Robustness of Monocular Depth Estimation with Procedural Scene Perturbations
- CoPa-SG: Dense Scene Graphs with Parametric and Proto-Relations
- Video Perception Models for 3D Scene Synthesis
- OptiScene: LLM-driven Indoor Scene Layout Generation via Scaled Human-aligned Data Synthesis and Multi-Stage Preference Optimization
- HOIverse: A Synthetic Scene Graph Dataset With Human Object Interactions
- Part2GS: Part-aware Modeling of Articulated Objects using 3D Gaussian Splatting
- DualTHOR: A Dual-Arm Humanoid Simulation Platform for Contingency-Aware Planning
- Efficient and Generalizable Environmental Understanding for Visual Navigation
- SpatialLM: Training Large Language Models for Structured Indoor Modeling
- Eye, Robot: Learning to Look to Act with a BC-RL Perception-Action Loop
- GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation
- LEO-VL: Efficient Scene Representation for Scalable 3D Vision-Language Learning
- Direct Numerical Layout Generation for 3D Indoor Scene Synthesis via Spatial Reasoning
- DEGround: An Effective Baseline for Ego-centric 3D Visual Grounding with a Homogeneous Framework
- Steerable Scene Generation with Post Training and Inference-Time Search
- Spatial Understanding from Videos: Structured Prompts Meet Simulation Data
- SEMNAV: A Semantic Segmentation-Driven Approach to Visual Semantic Navigation
- Growing Through Experience: Scaling Episodic Grounding in Language Models
- ArtiScene: Language-Driven Artistic 3D Scene Generation Through Image Intermediary
- Mobi-π: Mobilizing Your Robot Learning Policy
- LabUtopia: High-Fidelity Simulation and Hierarchical Benchmark for Scientific Embodied Agents
- DORAEMON: Decentralized Ontology-aware Reliable Agent with Enhanced Memory Oriented Navigation
- MORE: Mobile Manipulation Rearrangement Through Grounded Language Reasoning
- Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation
- BeliefMapNav: 3D Voxel-Based Belief Map for Zero-Shot Object Navigation
- GLEAM: Learning Generalizable Exploration Policy for Active Mapping in Complex 3D Indoor Scenes
- SD-OVON: A Semantics-aware Dataset and Benchmark Generation Pipeline for Open-Vocabulary Object Navigation in Dynamic Scenes
- Is Single-View Mesh Reconstruction Ready for Robotics?
- Beyond Needle(s) in the Embodied Haystack: Environment, Architecture, and Training Considerations for Long Context Reasoning
- MOON: Multi-Objective Optimization-Driven Object-Goal Navigation Using a Variable-Horizon Set-Orienteering Planner
- Long-Horizon Embodied Decision-Making via Multimodal Memory Compression
- When Replanning Becomes the Bottleneck: Budgeted Replanning for Embodied Agents
- SayCoNav: Utilizing Large Language Models for Adaptive Collaboration in Decentralized Multi-Robot Navigation
- LLM-BABYBENCH: Understanding and Evaluating Grounded Planning and Reasoning in LLMs
- Procedural Generation of Articulated Simulation-Ready Assets
- Towards Autonomous UAV Visual Object Search in City Space: Benchmark and Agentic Methodology
- Learning Semantic Priorities for Autonomous Target Search
- Towards Autonomous Micromobility through Scalable Urban Simulation
- A Survey of Robotic Navigation and Manipulation with Physics Simulators in the Era of Embodied AI
- Syn4D: A Multiview Synthetic 4D Dataset
- Predictive Spatio-Temporal Scene Graphs for Semi-Static Scenes
- CasaGPT: Cuboid Arrangement and Scene Assembly for Interior Design
- 3D-Belief: Embodied Belief Inference via Generative 3D World Modeling
- 3D Scene Graphs: Open Challenges and Future Directions
- Contact-Anchored Policies: Contact Conditioning Creates Strong Robot Utility Models
- Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding
- RoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot Learning
- Collaborating Action by Action: A Multi-agent LLM Framework for Embodied Reasoning
- What Makes Good Synthetic Training Data for Zero-Shot Stereo Matching?
- ForesightNav: Learning Scene Imagination for Efficient Exploration
- Multimodal Perception for Goal-oriented Navigation: A Survey
- DRAWER: Digital Reconstruction and Articulation With Environment Realism
- Digital Twin Generation from Visual Data: A Survey
- To See a World in a Living Context: Unified Indoor-Outdoor Urban World Generation
- iARCS: Iterative Agentic RL for Controllable 3D Scene Generation
- GenTe: Generative Real-world Terrains for General Legged Robot Locomotion Control
Discussions
Related