Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
2023/11/30 by Kristen Grauman, Andrew Westbury, Grauman, Kristen +197 · 92 citations
Computer Science · #Human Pose and Action Recognition #Multimodal Machine Learning Applications #Anomaly Detection Techniques and Applications
paper · pdf · doi:10.48550/arxiv.2311.18259
Abstract
We present Ego-Exo4D, a diverse, large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g., sports, music, dance, bike repair). 740 participants from 13 cities worldwide performed these activities in 123 different natural scene contexts, yielding long-form captures from 1 to 42 minutes each and 1,286 hours of video combined. The multimodal nature of the dataset is unprecedented: the video is accompanied by multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions -- including a novel "expert commentary" done by coaches and teachers and tailored to the skilled-activity domain. To push the frontier of first-person video understanding of skilled human activity, we also present a suite of benchmark tasks and their annotations, including fine-grained activity understanding, proficiency estimation, cross-view translation, and 3D hand/body pose. All resources are open sourced to fuel new research in the community. Project page: http://ego-exo4d-data.org/
Cited by
- Being-H0.7: A Latent World-Action Model from Egocentric Videos
- Human Motion Estimation with Everyday Wearables
- Mitty: Diffusion-based Human-to-Robot Video Generation
- OPENTOUCH: Bringing Full-Hand Touch to Real-World Interaction
- Prime and Reach: Synthesising Body Motion for Gaze-Primed Object Reach
- InfoTok: Adaptive Discrete Video Tokenizer via Information-Theoretic Compression
- Ego-EXTRA: video-language Egocentric Dataset for EXpert-TRAinee assistance
- SneakPeek: Future-Guided Instructional Streaming Video Generation
- Audio-Visual Camera Pose Estimation with Passive Scene Sounds and In-the-Wild Video
- The N-Body Problem: Parallel Execution from Single-Person Egocentric Video
- MultiEgo: A Multi-View Egocentric Video Dataset for 4D Scene Reconstruction
- An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
- VL-JEPA: Joint Embedding Predictive Architecture for Vision-language
- EgoX: Egocentric Video Generation from a Single Exocentric Video
- EgoCampus: Egocentric Pedestrian Eye Gaze Model and Dataset
- A Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and Reasoning
- Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?
- Towards Cross-View Point Correspondence in Vision-Language Models
- ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos
- ViDiC: Video Difference Captioning
- DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling
- PAI-Bench: A Comprehensive Benchmark For Physical AI
- Seeing without Pixels: Perception from Camera Trajectories
- V2-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence
- Exo2EgoSyn: Unlocking Foundation Video Generation Models for Exocentric-to-Egocentric Video Synthesis
- Distilling Counterfactual Reasoning from Language to Vision: Causal Graph Guided Post-Training for Video Understanding
- SkillSight: Efficient First-Person Skill Assessment with Gaze
- Robust Long-term Test-Time Adaptation for 3D Human Pose Estimation through Motion Discretization
- Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents
- Gaze Beyond the Frame: Forecasting Egocentric 3D Visual Span
- SFHand: A Streaming Framework for Language-guided 3D Hand Forecasting and Embodied Manipulation
- The Potential and Limitations of Vision-Language Models for Human Motion Understanding: A Case Study in Data-Driven Stroke Rehabilitation
- SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding
- Dexterity from Smart Lenses: Multi-Fingered Robot Manipulation with In-the-Wild Human Demonstrations
- Solving Spatial Supersensing Without Spatial Supersensing
- In-N-On: Scaling Egocentric Manipulation with in-the-wild and on-task Data
- RocSync: Millisecond-Accurate Temporal Synchronization for Heterogeneous Camera Systems
- Learning Skill-Attributes for Transferable Assessment in Video
- Building Egocentric Procedural AI Assistant: Methods, Benchmarks, and Challenges
- CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models
- EgoEMS: A High-Fidelity Multimodal Egocentric Dataset for Cognitive Assistance in Emergency Medical Services
- ViPRA: Video Prediction for Robot Actions
- SportR: A Benchmark for Multimodal Large Language Model Reasoning in Sports
- HAGI++: Head-Assisted Gaze Imputation and Generation
- A Step Toward World Models: A Survey on Robotic Manipulation
- The Quest for Generalizable Motion Generation: Data, Model, and Evaluation
- EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
- HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
- EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding
- TeleEgo: Benchmarking Egocentric AI Assistants in the Wild
- EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
- egoEMOTION: Egocentric Vision and Physiological Signals for Emotion and Personality Recognition in Real-World Tasks
- Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
- Vision-Based Mistake Analysis in Procedural Activities: A Review of Advances and Challenges
- Advances in 4D Representation: Geometry, Motion, and Interaction
- EventFormer: A Node-graph Hierarchical Attention Transformer for Action-centric Video Event Prediction
- Aria Gen 2 Pilot Dataset
- Eyes Wide Open: Ego Proactive Video-LLM for Streaming Video
- DeDelayed: Deleting Remote Inference Delay via On-Device Correction
- Learning Human Motion with Temporally Conditional Mamba
- Learning to Recognize Correctly Completed Procedure Steps in Egocentric Assembly Videos through Spatio-Temporal Modeling
- Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models
- TC-LoRA: Temporally Modulated Conditional LoRA for Adaptive Diffusion Control
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- Ego-Exo 3D Hand Tracking in the Wild with a Mobile Multi-Camera Rig
- EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark
- Learning Egocentric In-Hand Object Segmentation through Weak Supervision from Human Narrations
- EgoInstruct: An Egocentric Video Dataset of Face-to-face Instructional Interactions with Multi-modal LLM Benchmarking
- VLBiMan: Vision-Language Anchored One-Shot Demonstration Enables Generalizable Bimanual Robotic Manipulation
- Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer
- DanceEditor: Towards Iterative Editable Music-driven Dance Generation with Open-Vocabulary Descriptions
- Open-ended Hierarchical Streaming Video Understanding with Vision Language Models
- OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
- In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting
- Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation
- ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly
- Planning with Reasoning using Vision Language World Model
- Survey of Vision-Language-Action Models for Embodied Manipulation
- HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes
- Precise Action-to-Video Generation Through Visual Action Prompts
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
- Deformation Driven Suction Cups: A Mechanics-Based Approach to Wearable Electronics
- EgoMusic-driven Human Dance Motion Estimation with Skeleton Mamba
- AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
- Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions
- DOMR: Establishing Cross-View Segmentation via Dense Object Matching
- UniEgoMotion: A Unified Model for Egocentric Motion Reconstruction, Forecasting, and Generation
- Bidirectional Action Sequence Learning for Long-term Action Anticipation with Large Language Models
- H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation
- Reconstructing 4D Spatial Intelligence: A Survey
- EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
Related