vix.ing · top · new · best · stats

Action Scene Graphs for Long-Form Understanding of Egocentric Videos

2023/12/06 by Ivan Rodin, Rodin, Ivan, Antonino Furnari +7 · 21 citations
Computer Science · #Action (physics) #Anticipation (artificial intelligence) #Artificial intelligence #Automatic summarization #Computer Vision and Pattern Recognition (cs.CV) #Computer science #FOS: Computer and information sciences #Human Pose and Action Recognition #Human–computer interaction #Multimodal Machine Learning Applications #Natural language processing #Representation (politics) #Task (project management)

paper · pdf · doi:10.48550/arxiv.2312.03391

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2023/12/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We present Egocentric Action Scene Graphs (EASGs), a new representation for long-form understanding of egocentric videos. EASGs extend standard manually-annotated representations of egocentric videos, such as verb-noun action labels, by providing a temporally evolving graph-based description of the actions performed by the camera wearer, including interacted objects, their relationships, and how actions unfold in time. Through a novel annotation procedure, we extend the Ego4D dataset by adding manually labeled Egocentric Action Scene Graphs offering a rich set of annotations designed for long-from egocentric video understanding. We hence define the EASG generation task and provide a baseline approach, establishing preliminary benchmarks. Experiments on two downstream tasks, egocentric action anticipation and egocentric activity summarization, highlight the effectiveness of EASGs for long-form egocentric video understanding. We will release the dataset and the code to replicate experiments and annotations.

Cited by

Related