vix.ing · top · new · best · stats · spec

Video-based Human-Object Interaction Detection from Tubelet Tokens

2022/06/04 by Danyang Tu, Wei Sun, Tu, Danyang +7 · 2 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications #Video Surveillance and Tracking Methods

paper · pdf · doi:10.48550/arxiv.2206.01908

openalex publication_date 2022/06/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We present a novel vision Transformer, named TUTOR, which is able to learn tubelet tokens, served as highly-abstracted spatiotemporal representations, for video-based human-object interaction (V-HOI) detection. The tubelet tokens structurize videos by agglomerating and linking semantically-related patch tokens along spatial and temporal domains, which enjoy two benefits: 1) Compactness: each tubelet token is learned by a selective attention mechanism to reduce redundant spatial dependencies from others; 2) Expressiveness: each tubelet token is enabled to align with a semantic instance, i.e., an object or a human, across frames, thanks to agglomeration and linking. The effectiveness and efficiency of TUTOR are verified by extensive experiments. Results shows our method outperforms existing works by large margins, with a relative mAP gain of 16.14% on VidHOI and a 2 points gain on CAD-120 as well as a 4 × speedup.

Cited by

Related