Building Egocentric Procedural AI Assistant: Methods, Benchmarks, and Challenges
2025/11/17 by Junlong Li, Huaiyuan Xu, Li, Junlong +11
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2511.13261
openalex publication_date 2025/11/17 · openalex created_date 2025/11/19 · openalex updated_date 2026/07/28
Abstract
Driven by recent advances in vision-language models (VLMs) and egocentric perception research, the emerging topic of an egocentric procedural AI assistant (EgoProceAssist) is introduced to step-by-step support daily procedural tasks in a first-person view. In this paper, we start by identifying three core tasks in EgoProceAssist: egocentric procedural error detection, egocentric procedural learning, and egocentric procedural question answering, then introduce two enabling dimensions: real-time and streaming video understanding, and proactive interaction in procedural contexts. We define these tasks within a new taxonomy as the EgoProceAssist's essential functions and illustrate how they can be deployed in real-world scenarios for daily activity assistants. Specifically, our work encompasses a comprehensive review of current techniques, relevant datasets, and evaluation metrics across these five core areas. To clarify the gap between the proposed EgoProceAssist and existing VLM-based assistants, we conduct novel experiments to provide a comprehensive evaluation of representative VLM-based methods. Through these findings and our technical analysis, we discuss the challenges ahead and suggest future research directions. Furthermore, an exhaustive list of this study is publicly available in an active repository that continuously collects the latest work: https://github.com/z1oong/Building-Egocentric-Procedural-AI-Assistant.
Citations
- StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
- EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning
- EgoSocial: Benchmarking Proactive Intervention Ability of Omnimodal LLMs via Egocentric Social Interaction Perception
- Inference-stage Adaptation-projection Strategy Adapts Diffusion Policy to Cross-manipulators Scenarios
- Planning with Reasoning using Vision Language World Model
- CoopTrack: Exploring End-to-End Learning for Efficient Cooperative Sequential Perception
- Procedure Learning via Regularized Gromov-Wasserstein Optimal Transport
- ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models
- Egocentric Human-Object Interaction Detection: A New Benchmark and Method
- Ming-Omni: A Unified Multimodal Model for Perception and Generation
- From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge
- Technical Report for Egocentric Mistake Detection for the HoloAssist Challenge
- Proactive Assistant Dialogue Generation from Streaming Egocentric Videos
- EgoVLM: Policy Optimization for Egocentric Video Understanding
- Towards Comprehensive Scene Understanding: Integrating First and Third-Person Views for LVLMs
- HCQA-1.5 @ Ego4D EgoSchema Challenge 2025
- RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language
- HiERO: understanding the hierarchy of human behavior enhances reasoning on egocentric videos
- StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming Assistant
- How Do I Do That? Synthesizing 3D Hand Motion and Contacts for Everyday Interactions
- ANNEXE: Unified Analyzing, Answering, and Pixel Grounding for Egocentric Interaction
- Modeling Multiple Normal Action Representations for Error Detection in Procedural Tasks
- Qwen2.5-Omni Technical Report
- Challenges and Trends in Egocentric Vision: A Survey
- Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos
- EgoBlind: Towards Egocentric Visual Assistance for the Blind
- EgoLife: Towards Egocentric Life Assistant
- SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding
- EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
- EgoAgent: A Joint Predictive Agent Model in Egocentric Worlds
- Hier-EgoPack: Hierarchical Egocentric Video Understanding with Diverse Task Perspectives
- Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge
- YETI (YET to Intervene) Proactive Interventions by Multimodal AI Agents in Augmented Reality Tasks
- OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
- Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction
- VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
- Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding
- Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model
- Transparent and Coherent Procedural Mistake Detection
- InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
- SocialMind: LLM-based Proactive AR Social Assistive System with Human-like Perception for In-situ Live Interactions
- VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format
- IKEA Manuals at Work: 4D Grounding of Assembly Instructions on Internet Videos
- StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
- TI-PREGO: Chain of Thought and In-Context Learning for Online Mistake Detection in PRocedural EGOcentric Videos
- Egocentric and Exocentric Methods: A Short Survey
- GPT-4o System Card
- Proactive Agent: Shifting LLM Agents from Reactive Responses to Active Assistance
- Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities
- EgoOops: A Dataset for Mistake Action Detection from Egocentric Videos referring to Procedural Texts
- AMEGO: Active Memory from long EGOcentric videos
- VideoLLM-MoD: Efficient Video-Language Streaming with Mixture-of-Depths Vision Computation
- Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
- Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
- Egocentric Vision Language Planning
- VITA: Towards Open-Source Interactive Omni Multimodal LLM
- VideoQA in the Era of LLMs: An Empirical Study
- LLaVA-OneVision: Easy Visual Task Transfer
- CaRe-Ego: Contact-aware Relationship Modeling for Egocentric Interactive Hand-object Segmentation
- EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation
- HCQA @ Ego4D EgoSchema Challenge 2024
- AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding
- VideoLLM-online: Online Video Large Language Model for Streaming Video
- Gazing Into Missteps: Leveraging Eye-Gaze for Unsupervised Mistake Detection in Egocentric Videos of Skilled Human Activities
- Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
- VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
- Video Question Answering for People with Visual Impairments Using an Egocentric 360-Degree Camera
- Streaming Long Video Understanding with Large Language Models
- HOI-Ref: Hand-Object Interaction Referral in Egocentric Vision
- Spatial Cognition from Egocentric Video: Out of Sight, Not Out of Mind
- PREGO: online mistake detection in PRocedural EGOcentric videos
- InternLM2 Technical Report
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- CaptainCook4D: A Dataset for Understanding Errors in Procedural Activities
- Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos
- Grounded Question-Answering in Long Egocentric Videos
- LifelongMemory: Leveraging LLMs for Answering Queries in Long-form Egocentric Videos
- Action Scene Graphs for Long-Form Understanding of Egocentric Videos
- LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning
- Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
- Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
- Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
- United We Stand, Divided We Fall: UnityGraph for Unsupervised Procedure Learning from Videos
- IndustReal: A Dataset for Procedure Step Recognition Handling Execution Errors in Egocentric Videos in an Industrial-Like Setting
- HoloAssist: an Egocentric Human Interaction Dataset for Interactive AI Assistants in the Real World
- ProAgent: Building Proactive Cooperative Agents with Large Language Models
- EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding
- An Outlook into the Future of Egocentric Vision
- AntGPT: Can Large Language Models Help Long-term Action Anticipation from Videos?
- Every Mistake Counts in Assembly
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the Backbone
- Reading Between the Lanes: Text VideoQA on the Road
- Egocentric Planning for Scalable Embodied Task Achievement
- A Survey on Proactive Dialogue Systems: Problems, Methods, and Prospects
- I3D: Transformer architectures with input-dependent dynamic depth for speech recognition
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- NaQ: Leveraging Narrations as Queries to Supervise Episodic Memory
- InternVideo: General Video Foundation Models via Generative and Discriminative Learning
- InternVideo-Ego4D: A Pack of Champion Solutions to Ego4D Challenges
- A Simple Transformer-Based Model for Ego4D Natural Language Queries Challenge
- Temporal Action Segmentation: An Analysis of Modern Techniques
- Automatic Chain of Thought Prompting in Large Language Models
- Real-time Online Video Detection with Temporal Smoothing Transformers
- In the Eye of Transformer: Global-Local Correlation for Egocentric Gaze Estimation
- My View is the Best View: Procedure Learning from Egocentric Videos
- PP-OCRv3: More Attempts for the Improvement of Ultra Lightweight OCR System
- Egocentric Video-Language Pretraining
- Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities
- AssistQ: Affordance-centric Question-driven Task Completion for Egocentric Assistant
- Learning To Recognize Procedural Activities with Distant Supervision
- Ego4D: Around the World in 3,000 Hours of Egocentric Video
- Predicting the Future from First Person (Egocentric) Vision: A Survey
- OadTR: Online Action Detection with Transformers
- Prototypical Graph Contrastive Learning
- A Survey of Embodied AI: From Simulators to Research Tasks
- Is Space-Time Attention All You Need for Video Understanding?
- Just Ask: Learning to Answer Questions from Millions of Narrated Videos
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Deep Learning for Vision-based Prediction: A Survey
- The EPIC-KITCHENS Dataset: Collection, Challenges and Baselines
- IPN Hand: A Video Dataset and Benchmark for Real-Time Continuous Hand Gesture Recognition
- RAFT: Recurrent All-Pairs Field Transforms for Optical Flow
- Oops! Predicting Unintentional Action in Video
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language\n Generation, Translation, and Comprehension
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million\n Narrated Video Clips
- Unsupervised learning of action classes with continuous temporal embedding
- Cross-task weakly supervised learning from instructional videos
- COIN: A Large-scale Dataset for Comprehensive Instructional Video Analysis
- Class-Balanced Loss Based on Effective Number of Samples
- Representation Learning with Contrastive Predictive Coding
- Human Action Recognition and Prediction: A Survey
- Human Action Recognition and Prediction: A Survey
- NeuralNetwork-Viterbi: A Framework for Weakly Supervised Video Learning
- Scaling Egocentric Vision: The EPIC-KITCHENS Dataset
- VizWiz Grand Challenge: Answering Visual Questions from Blind People
- Real-world Anomaly Detection in Surveillance Videos
- Embodied Question Answering
- To prune, or not to prune: exploring the efficacy of pruning for model compression
- Localizing Moments in Video with Natural Language
- Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset
- Towards Automatic Learning of Procedures from Web Instructional Videos
- First-Person Activity Forecasting with Online Inverse Reinforcement\n Learning
- Semantic Scene Completion from a Single Depth Image
- node2vec: Scalable Feature Learning for Networks
- Online Action Detection
- Deep Residual Learning for Image Recognition
- Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
- Unsupervised Learning from Narrated Instruction Videos
- Unsupervised Semantic Parsing of Video Collections
- A Survey on Recent Advances of Computer Vision Algorithms for Egocentric\n Video
- STEPs: Self-Supervised Key Step Extraction and Localization from Unlabeled Procedural Videos
Related