Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
2023/11/30 by Kristen Grauman, Andrew Westbury, Grauman, Kristen +197 · 165 citations
Computer Science · #Human Pose and Action Recognition #Multimodal Machine Learning Applications #Anomaly Detection Techniques and Applications
paper · pdf · doi:10.48550/arxiv.2311.18259
Abstract
We present Ego-Exo4D, a diverse, large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g., sports, music, dance, bike repair). 740 participants from 13 cities worldwide performed these activities in 123 different natural scene contexts, yielding long-form captures from 1 to 42 minutes each and 1,286 hours of video combined. The multimodal nature of the dataset is unprecedented: the video is accompanied by multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions -- including a novel "expert commentary" done by coaches and teachers and tailored to the skilled-activity domain. To push the frontier of first-person video understanding of skilled human activity, we also present a suite of benchmark tasks and their annotations, including fine-grained activity understanding, proficiency estimation, cross-view translation, and 3D hand/body pose. All resources are open sourced to fuel new research in the community. Project page: http://ego-exo4d-data.org/
Cited by
- Being-H0.7: A Latent World-Action Model from Egocentric Videos
- Human Motion Estimation with Everyday Wearables
- Mitty: Diffusion-based Human-to-Robot Video Generation
- OPENTOUCH: Bringing Full-Hand Touch to Real-World Interaction
- Prime and Reach: Synthesising Body Motion for Gaze-Primed Object Reach
- InfoTok: Adaptive Discrete Video Tokenizer via Information-Theoretic Compression
- Ego-EXTRA: video-language Egocentric Dataset for EXpert-TRAinee assistance
- SneakPeek: Future-Guided Instructional Streaming Video Generation
- Audio-Visual Camera Pose Estimation with Passive Scene Sounds and In-the-Wild Video
- The N-Body Problem: Parallel Execution from Single-Person Egocentric Video
- MultiEgo: A Multi-View Egocentric Video Dataset for 4D Scene Reconstruction
- An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
- VL-JEPA: Joint Embedding Predictive Architecture for Vision-language
- EgoX: Egocentric Video Generation from a Single Exocentric Video
- EgoCampus: Egocentric Pedestrian Eye Gaze Model and Dataset
- A Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and Reasoning
- Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?
- Towards Cross-View Point Correspondence in Vision-Language Models
- ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos
- ViDiC: Video Difference Captioning
- DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling
- PAI-Bench: A Comprehensive Benchmark For Physical AI
- Seeing without Pixels: Perception from Camera Trajectories
- V2-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence
- Exo2EgoSyn: Unlocking Foundation Video Generation Models for Exocentric-to-Egocentric Video Synthesis
- Distilling Counterfactual Reasoning from Language to Vision: Causal Graph Guided Post-Training for Video Understanding
- SkillSight: Efficient First-Person Skill Assessment with Gaze
- Robust Long-term Test-Time Adaptation for 3D Human Pose Estimation through Motion Discretization
- Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents
- Gaze Beyond the Frame: Forecasting Egocentric 3D Visual Span
- SFHand: A Streaming Framework for Language-guided 3D Hand Forecasting and Embodied Manipulation
- The Potential and Limitations of Vision-Language Models for Human Motion Understanding: A Case Study in Data-Driven Stroke Rehabilitation
- SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding
- Dexterity from Smart Lenses: Multi-Fingered Robot Manipulation with In-the-Wild Human Demonstrations
- Solving Spatial Supersensing Without Spatial Supersensing
- In-N-On: Scaling Egocentric Manipulation with in-the-wild and on-task Data
- RocSync: Millisecond-Accurate Temporal Synchronization for Heterogeneous Camera Systems
- Learning Skill-Attributes for Transferable Assessment in Video
- Building Egocentric Procedural AI Assistant: Methods, Benchmarks, and Challenges
- CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models
- EgoEMS: A High-Fidelity Multimodal Egocentric Dataset for Cognitive Assistance in Emergency Medical Services
- ViPRA: Video Prediction for Robot Actions
- SportR: A Benchmark for Multimodal Large Language Model Reasoning in Sports
- HAGI++: Head-Assisted Gaze Imputation and Generation
- A Step Toward World Models: A Survey on Robotic Manipulation
- The Quest for Generalizable Motion Generation: Data, Model, and Evaluation
- EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
- HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
- EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding
- TeleEgo: Benchmarking Egocentric AI Assistants in the Wild
- EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
- egoEMOTION: Egocentric Vision and Physiological Signals for Emotion and Personality Recognition in Real-World Tasks
- Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
- Vision-Based Mistake Analysis in Procedural Activities: A Review of Advances and Challenges
- Advances in 4D Representation: Geometry, Motion, and Interaction
- EventFormer: A Node-graph Hierarchical Attention Transformer for Action-centric Video Event Prediction
- Aria Gen 2 Pilot Dataset
- Eyes Wide Open: Ego Proactive Video-LLM for Streaming Video
- DeDelayed: Deleting Remote Inference Delay via On-Device Correction
- Learning Human Motion with Temporally Conditional Mamba
- Learning to Recognize Correctly Completed Procedure Steps in Egocentric Assembly Videos through Spatio-Temporal Modeling
- Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models
- TC-LoRA: Temporally Modulated Conditional LoRA for Adaptive Diffusion Control
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- Ego-Exo 3D Hand Tracking in the Wild with a Mobile Multi-Camera Rig
- EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark
- Learning Egocentric In-Hand Object Segmentation through Weak Supervision from Human Narrations
- EgoInstruct: An Egocentric Video Dataset of Face-to-face Instructional Interactions with Multi-modal LLM Benchmarking
- VLBiMan: Vision-Language Anchored One-Shot Demonstration Enables Generalizable Bimanual Robotic Manipulation
- Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer
- DanceEditor: Towards Iterative Editable Music-driven Dance Generation with Open-Vocabulary Descriptions
- Open-ended Hierarchical Streaming Video Understanding with Vision Language Models
- OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
- In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting
- Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation
- ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly
- Planning with Reasoning using Vision Language World Model
- PlayerOne: Egocentric World Simulator
- Survey of Vision-Language-Action Models for Embodied Manipulation
- HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes
- Precise Action-to-Video Generation Through Visual Action Prompts
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
- Deformation Driven Suction Cups: A Mechanics-Based Approach to Wearable Electronics
- EgoMusic-driven Human Dance Motion Estimation with Skeleton Mamba
- AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
- Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions
- DOMR: Establishing Cross-View Segmentation via Dense Object Matching
- UniEgoMotion: A Unified Model for Egocentric Motion Reconstruction, Forecasting, and Generation
- Is Tracking really more challenging in First Person Egocentric Vision?
- Bidirectional Action Sequence Learning for Long-term Action Anticipation with Large Language Models
- H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation
- Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
- Reconstructing 4D Spatial Intelligence: A Survey
- EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
- MobileEgo Anywhere: Open Infrastructure for long horizon egocentric data on commodity hardware
- Scene Text Detection and Recognition "in light of" Challenging Environmental Conditions using Aria Glasses Egocentric Vision Cameras
- EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos
- CuriosAI Submission to the EgoExo4D Proficiency Estimation Challenge 2025
- COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention
- THU-Warwick Submission for EPIC-KITCHEN Challenge 2025: Semi-Supervised Video Object Segmentation
- UFM: A Simple Path towards Unified Dense Correspondence with Flow
- A Survey: Learning Embodied Intelligence from Physical Simulators and World Models
- RGC-VQA: An Exploration Database for Robotic-Generated Video Quality Assessment
- EgoM2P: Egocentric Multimodal Multitask Pretraining
- Embodied AI Agents: Modeling the World
- EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception
- Whole-Body Conditioned Egocentric Video Prediction
- DemoDiffusion: One-Shot Human Imitation using pre-trained Diffusion Policy
- Systematic Comparison of Projection Methods for Monocular 3D Human Pose Estimation on Fisheye Images
- VisualChef: Generating Visual Aids in Cooking via Mask Inpainting
- 4DGT: Learning a 4D Gaussian Transformer Using Real-World Monocular Videos
- EgoWorld: Translating Exocentric View to Egocentric View using Rich Exocentric Observations
- Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
- MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks
- Foundation Models in Autonomous Driving: A Survey on Scenario Generation and Scenario Analysis
- Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision
- O-MaMa: Learning Object Mask Matching between Egocentric and Exocentric Views
- ExAct: A Video-Language Benchmark for Expert Action Analysis
- Proactive Assistant Dialogue Generation from Streaming Egocentric Videos
- Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025
- WorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural Planning
- Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision
- Seeing in the Dark: Benchmarking Egocentric 3D Vision with the Oxford Day-and-Night Dataset
- EgoBrain: Synergizing Minds and Eyes For Human Action Understanding
- Improving Keystep Recognition in Ego-Video via Dexterous Focus
- Keystep Recognition using Graph Neural Networks
- Sequence-Based Identification of First-Person Camera Wearers in Third-Person Views
- Reducing Annotation Burden in Physical Activity Research Using Vision-Language Models
- PCIEPose Solution for EgoExo4D Pose and Proficiency Estimation Challenge
- Leadership Assessment in Pediatric Intensive Care Unit Team Training
- EgoExOR: An Ego-Exo-Centric Operating Room Dataset for Surgical Activity Understanding
- Reading Recognition in the Wild
- Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks
- Towards Comprehensive Scene Understanding: Integrating First and Third-Person Views for LVLMs
- Predicting Implicit Arguments in Procedural Video Instructions
- OpenHOI: Open-World Hand-Object Interaction Synthesis with Multimodal Large Language Model
- Exploring the Limits of Vision-Language-Action Manipulations in Cross-task Generalization
- Egocentric Action-aware Inertial Localization in Point Clouds with Vision-Language Guidance
- SkillFormer: Unified Multi-View Video Understanding for Proficiency Estimation
- DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams
- EgoIntent: A Pre-Outcome Micro-Step Benchmark for Understanding What, Why, and Next
- SITE: towards Spatial Intelligence Thorough Evaluation
- AdaDINO: Context-Adaptive DINO-Distilled Vision Foundation Models for Efficient Open-Vocabulary Edge Inference
- Grounding Task Assistance with Multimodal Cues from a Single Demonstration
- Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models
- Learning Streaming Video Representation via Multitask Training
- Action100M: A Large-scale Video Action Dataset
- AIDE: Automated Instruction via Distilled Expertise for Reference-Free Motor Skill Coaching
- Dynamic Camera Poses and Where to Find Them
- EgoCHARM: Resource-Efficient Hierarchical Activity Recognition using an Egocentric IMU Sensor
- Hierarchical and Multimodal Data for Daily Activity Understanding
- RoboReact: Agentic Skill Distillation from Generated Egocentric Videos for Generalizable Whole-Body Manipulation
- SiMDex: Mining Similar Egocentric Videos for Cross-Embodiment Dexterous Manipulation
- InstructionBench: An Instructional Video Understanding Benchmark
- ProbRes: Probabilistic Jump Diffusion for Open-World Egocentric Activity Recognition
- Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
- Towards Understanding Camera Motions in Any Video
- EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
- LAWM-3D: Learning 3D-Aware Latent Actions from Human Videos for Generalizable Robot World Models
- JoyAI-RA 0.5: Scaling Robot Manipulation Learning via Dual Action Alignment
- VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances
- Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
- MAPLE: Encoding Dexterous Robotic Manipulation Priors Learned From Egocentric Videos
- Unsupervised Ego- and Exo-centric Dense Procedural Activity Captioning via Gaze Consensus Adaptation
Related