Ego4D: Around the World in 3,000 Hours of Egocentric Video
2021/10/13 by Kristen Grauman, Andrew Westbury, Grauman, Kristen +166 · 207 citations
Computer Science · Social Sciences · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Participatory Visual Research Methods #Video Surveillance and Tracking Methods
paper · doi:10.48550/arxiv.2110.07058
openalex publication_date 2021/10/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countries. The approach to collection is designed to uphold rigorous privacy and ethics standards with consenting participants and robust de-identification procedures where relevant. Ego4D dramatically expands the volume of diverse egocentric video footage publicly available to the research community. Portions of the video are accompanied by audio, 3D meshes of the environment, eye gaze, stereo, and/or synchronized videos from multiple egocentric cameras at the same event. Furthermore, we present a host of new benchmark challenges centered around understanding the first-person visual experience in the past (querying an episodic memory), present (analyzing hand-object manipulation, audio-visual conversation, and social interactions), and future (forecasting activities). By publicly sharing this massive annotated dataset and benchmark suite, we aim to push the frontier of first-person perception. Project page: https://ego4d-data.org/
Cited by
- Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
- Emergence of Human to Robot Transfer in Vision-Language-Action Models
- TimePLE: Rethinking Temporal Representation for Video Temporal Grounding
- Being-H0.7: A Latent World-Action Model from Egocentric Videos
- UniTacHand: Unified Spatio-Tactile Representation for Human to Robotic Hand Skill Transfer
- Human Motion Estimation with Everyday Wearables
- Learning Skills from Action-Free Videos
- QuantiPhy: A Quantitative Benchmark Evaluating Physical Reasoning Abilities of Vision-Language Models
- PEDESTRIAN: An Egocentric Vision Dataset for Obstacle Detection on Pavements
- Revisiting the Learning Objectives of Vision-Language Reward Models
- Mitty: Diffusion-based Human-to-Robot Video Generation
- OPENTOUCH: Bringing Full-Hand Touch to Real-World Interaction
- GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA Manipulation
- PhysBrain: Human Egocentric Data as a Bridge from Vision Language Models to Physical Intelligence
- Prime and Reach: Synthesising Body Motion for Gaze-Primed Object Reach
- Sceniris: A Fast Procedural Scene Generation Framework
- See It Before You Grab It: Deep Learning-based Action Anticipation in Basketball
- Seeing is Believing (and Predicting): Context-Aware Multi-Human Behavior Prediction with Vision Language Models
- Large Video Planner Enables Generalizable Robot Control
- GateFusion: Hierarchical Gated Cross-Modal Fusion for Active Speaker Detection
- TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
- Elastic3D: Controllable Stereo Video Conversion with Guided Latent Decoding
- KFS-Bench: Comprehensive Evaluation of Key Frame Sampling in Long Video Understanding
- World Models Can Leverage Human Videos for Dexterous Manipulation
- Adapting MLLMs for Nuanced Video Retrieval
- Moment and Highlight Detection via MLLM Frame Segmentation
- The N-Body Problem: Parallel Execution from Single-Person Egocentric Video
- Benchmarking the Generality of Vision-Language-Action Models
- An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
- UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Models
- VL-JEPA: Joint Embedding Predictive Architecture for Vision-language
- BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models
- EgoX: Egocentric Video Generation from a Single Exocentric Video
- VLD: Visual Language Goal Distance for Reinforcement Learning Navigation
- EgoCampus: Egocentric Pedestrian Eye Gaze Model and Dataset
- A Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and Reasoning
- See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations
- Know-Show: Benchmarking Video-Language Models on Spatio-Temporal Grounded Reasoning
- Hoi! - A Multimodal Dataset for Force-Grounded, Cross-View Articulated Manipulation
- Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?
- COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence
- PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement
- EgoLCD: Egocentric Video Generation with Long Context Diffusion
- StreamEQA: Towards Streaming Video Understanding for Embodied Scenarios
- FALCON: Actively Decoupled Visuomotor Policies for Loco-Manipulation with Foundation-Model-Based Coordination
- Generalized Event Partonomy Inference with Structured Hierarchical Predictive Learning
- Unique Lives, Shared World: Learning from Single-Life Videos
- ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos
- EEA: Exploration-Exploitation Agent for Long Video Understanding
- KM-ViPE: Online Tightly Coupled Vision-Language-Geometry Fusion for Open-Vocabulary Semantic SLAM
- StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
- See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models
- HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics
- Audio-Visual World Models: Grounding Multisensory Imagination for Embodied Agents
- SocialFusion: Addressing Social Degradation in Pre-trained Vision-Language Models
- Video-CoM: Interactive Video Reasoning via Chain of Manipulations
- Learning to Refuse: Refusal-Aware Reinforcement Fine-Tuning for Hard-Irrelevant Queries in Video Temporal Grounding
- Geometrically-Constrained Agent for Spatial Reasoning
- Seeing without Pixels: Perception from Camera Trajectories
- From Observation to Action: Latent Action-based Primitive Segmentation for VLA Pre-training in Industrial Settings
- Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric Videos
- Unifying Perception and Action: A Hybrid-Modality Pipeline with Implicit Visual Chain-of-Thought for Robotic Action Generation
- Gaze Beyond the Frame: Forecasting Egocentric 3D Visual Span
- TRANSPORTER: Transferring Visual Semantics from VLM Manifolds
- Sequence-Adaptive Video Prediction in Continuous Streams using Diffusion Noise Optimization
- EgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses
- SFHand: A Streaming Framework for Language-guided 3D Hand Forecasting and Embodied Manipulation
- Multi-speaker Attention Alignment for Multimodal Social Interaction
- The Potential and Limitations of Vision-Language Models for Human Motion Understanding: A Case Study in Data-Driven Stroke Rehabilitation
- Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions
- OmniGround: A Comprehensive Spatio-Temporal Grounding Benchmark for Real-World Complex Scenarios
- Dexterity from Smart Lenses: Multi-Fingered Robot Manipulation with In-the-Wild Human Demonstrations
- TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
- Click2Graph: Interactive Panoptic Video Scene Graphs from a Single Click
- In-N-On: Scaling Egocentric Manipulation with in-the-wild and on-task Data
- Uni-Hand: Universal Hand Motion Forecasting in Egocentric Views
- Building Egocentric Procedural AI Assistant: Methods, Benchmarks, and Challenges
- EgoCogNav: Cognition-aware Human Egocentric Navigation
- Attentive Feature Aggregation or: How Policies Learn to Stop Worrying about Robustness and Attend to Task-Relevant Visual Cues
- EgoEMS: A High-Fidelity Multimodal Egocentric Dataset for Cognitive Assistance in Emergency Medical Services
- ViPRA: Video Prediction for Robot Actions
- Countering Multi-modal Representation Collapse through Rank-targeted Fusion
- LiveStar: Live Streaming Assistant for Real-World Online Video Understanding
- Let Me Show You: Learning by Retrieving from Egocentric Video for Robotic Manipulation
- iFlyBot-VLM Technical Report
- Visual Spatial Tuning
- Tracking and Understanding Object Transformations
- Cambrian-S: Towards Spatial Supersensing in Video
- Temporal Zoom Networks: Distance Regression and Continuous Depth for Efficient Action Localization
- NVIDIA Nemotron Nano V2 VL
- Accelerating Physical Property Reasoning for Augmented Visual Cognition
- XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
- The Pervasive Blind Spot: Benchmarking VLM Inference Risks on Everyday Personal Videos
- iFlyBot-VLA Technical Report
- Can MLLMs Read the Room? A Multimodal Benchmark for Verifying Truthfulness in Multi-Party Social Interactions
- A Step Toward World Models: A Survey on Robotic Manipulation
- Face De-Identification: A Domain-Centric Survey from Capture to Processing
- HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
- CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Games
- CG-World: A Large-Scale World-State Dataset and Protocol for World Models
- From Passive Video to Editable Experience: Physically Grounded Experience Synthesis for Embodied Intelligence
- EDVD-LLaMA: Explainable Deepfake Video Detection via Multimodal Large Language Model Reasoning
- TeleEgo: Benchmarking Egocentric AI Assistants in the Wild
- EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
- A Cocktail-Party Benchmark: Multi-Modal dataset and Comparative Evaluation Results
- HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling
- Look and Tell: A Dataset for Multimodal Grounding Across Egocentric and Exocentric Views
- Benchmarking Egocentric Multimodal Goal Inference for Assistive Wearable Agents
- egoEMOTION: Egocentric Vision and Physiological Signals for Emotion and Personality Recognition in Real-World Tasks
- Gaze-VLM:Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding
- Towards Physics-informed Spatial Intelligence with Human Priors: An Autonomous Driving Pilot Study
- Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
- EmbodiedBrain: Expanding Performance Boundaries of Task Planning for Embodied Intelligence
- DMC3: Dual-Modal Counterfactual Contrastive Construction for Egocentric Video Question Answering
- Vision-Based Mistake Analysis in Procedural Activities: A Review of Advances and Challenges
- Advances in 4D Representation: Geometry, Motion, and Interaction
- X-Ego: Acquiring Team-Level Tactical Situational Awareness via Cross-Egocentric Contrastive Video Representation Learning
- Is This Tracker On? A Benchmark Protocol for Dynamic Tracking
- Event-Grounding Graph: Unified Spatio-Temporal Scene Graph from Robotic Observations
- AI-Boosted Video Annotation: Assessing the Process Enhancement
- LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
- Training-free Online Video Step Grounding
- EventFormer: A Node-graph Hierarchical Attention Transformer for Action-centric Video Event Prediction
- TOUCH: Text-guided Controllable Generation of Free-Form Hand-Object Interactions
- Eyes Wide Open: Ego Proactive Video-LLM for Streaming Video
- DeDelayed: Deleting Remote Inference Delay via On-Device Correction
- Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs
- EgoSocial: Benchmarking Proactive Intervention Ability of Omnimodal LLMs via Egocentric Social Interaction Perception
- VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
- Learning to Recognize Correctly Completed Procedure Steps in Egocentric Assembly Videos through Spatio-Temporal Modeling
- mmWalk: Towards Multi-modal Multi-view Walking Assistance
- Seeing My Future: Predicting Situated Interaction Behavior in Virtual Reality
- UniCoD: Enhancing Robot Policy via Unified Continuous and Discrete Representation Learning
- ESCA: Contextualizing Embodied Agents via Scene-Graph Generation
- Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models
- Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding
- DexMan: Learning Bimanual Dexterous Manipulation from Human and Generated Videos
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- Addressing the ID-Matching Challenge in Long Video Captioning
- EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark
- ActiveUMI: Robotic Manipulation with Active Perception from Robot-Free Human Demonstrations
- EmbodiSwap for Zero-Shot Robot Imitation Learning
- AdaRD-key: Adaptive Relevance-Diversity Keyframe Sampling for Long-form Video understanding
- EgoTraj-Bench: Towards Robust Trajectory Prediction Under Ego-view Noisy Observations
- TSalV360: A Method and Dataset for Text-driven Saliency Detection in 360-Degrees Videos
- Learning Egocentric In-Hand Object Segmentation through Weak Supervision from Human Narrations
- Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play
- NeMo: Needle in a Montage for Video-Language Understanding
- Mash, Spread, Slice! Learning to Manipulate Object States via Visual Spatial Progress
- MimicDreamer: Aligning Human and Robot Demonstrations for Scalable VLA Training
- EgoInstruct: An Egocentric Video Dataset of Face-to-face Instructional Interactions with Multi-modal LLM Benchmarking
- Developing Vision-Language-Action Model from Egocentric Videos
- Parse-Augment-Distill: Learning Generalizable Bimanual Visuomotor Policies from Single Human Video
- Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer
- VIR-Bench: Evaluating Geospatial and Temporal Understanding of MLLMs via Travel Video Itinerary Reconstruction
- Latent Action Pretraining Through World Modeling
- Person Identification from Egocentric Human-Object Interactions using 3D Hand Pose
- Segment-to-Act: Label-Noise-Robust Action-Prompted Video Segmentation Towards Embodied Intelligence
- Conversational Orientation Reasoning: Egocentric-to-Allocentric Navigation with Multimodal Chain-of-Thought
- RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
- MICA: Multi-Agent Industrial Coordination Assistant
- 4D Visual Pre-training for Robot Learning
- Using LLMs for Late Multimodal Sensor Fusion for Activity Recognition
- Video Understanding by Design: How Datasets Shape Architectures and Insights
- Diffusion-Based Action Recognition Generalizes to Untrained Domains
- Chirality in Action: Time-Aware Video Representation Learning by Latent Straightening
- In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting
- Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation
- ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly
- Planning with Reasoning using Vision Language World Model
- Robix: A Unified Model for Robot Interaction, Reasoning and Planning
- HERO-VQL: Hierarchical, Egocentric and Robust Visual Query Localization
- Representation Learning with Adaptive Superpixel Coding
- StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
- Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models
- MovieCORE: COgnitive REasoning in Movies
- Rethinking Human-Object Interaction Evaluation for both Vision-Language Models and HOI-Specific Methods
- Survey of Vision-Language-Action Models for Embodied Manipulation
- LookOut: Real-World Humanoid Egocentric Navigation
- HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes
- Precise Action-to-Video Generation Through Visual Action Prompts
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
- QuickMerge++: Fast Token Merging with Autoregressive Prior
- Deformation Driven Suction Cups: A Mechanics-Based Approach to Wearable Electronics
- Generating Dialogues from Egocentric Instructional Videos for Task Assistance: Dataset, Method and Benchmark
- Ensembling Synchronisation-based and Face-Voice Association Paradigms for Robust Active Speaker Detection in Egocentric Recordings
- Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
- Enhancing Monocular 3D Hand Reconstruction with Learned Texture Priors
- Masquerade: Learning from In-the-wild Human Videos using Data-Editing
- Boosting Action-Information via a Variational Bottleneck on Unlabelled Robot Videos
- TAR-TVG: Enhancing VLMs with Timestamp Anchor-Constrained Reasoning for Temporal Video Grounding
- AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical Reasoning
- Enhancing Egocentric Object Detection in Static Environments using Graph-based Spatial Anomaly Detection and Correction
- SafePLUG: Empowering Multimodal LLMs with Pixel-Level Insight and Temporal Grounding for Traffic Accident Understanding
- A Survey on Video Temporal Grounding with Multimodal Large Language Model
- Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions
- OpenLifelogQA: An Open-Ended Multi-Modal Lifelog Question-Answering Dataset
- VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
- StreamAgent: Towards Anticipatory Agents for Streaming Video Understanding
- EgoTrigger: Toward Audio-Driven Image Capture for Human Memory Enhancement in All-Day Energy-Efficient Smart Glasses
- UniEgoMotion: A Unified Model for Egocentric Motion Reconstruction, Forecasting, and Generation
- Fine-grained Spatiotemporal Grounding on Egocentric Videos
- iSafetyBench: A video-language benchmark for safety in industrial environment
- Bidirectional Action Sequence Learning for Long-term Action Anticipation with Large Language Models
- RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping
- villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models
Related