2014/10/21 by Ran Xu, Gang Chen, Xu, Ran +7 · 1 citation
Computer Science · Engineering · Mathematics · #Action (physics) #Action recognition #Artificial intelligence #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Computer vision #Discriminative model #FOS: Computer and information sciences #Focus (optics) #Gait Recognition and Analysis #Human Pose and Action Recognition #Human motion #Mathematics #Motion (physics) #Motion capture #Multimodal Machine Learning Applications #Ranging #Representation (politics) #Robot #Robotics #Space (punctuation) #Variable (mathematics) #cs.CV
paper · pdf · doi:10.48550/arxiv.1410.5861
13 pages
arxiv created 2014/10/21 · openalex publication_date 2014/10/21 · arxiv updated 2014/10/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/06
The focus of the action understanding literature has predominately been classification, how- ever, there are many applications demanding richer action understanding such as mobile robotics and video search, with solutions to classification, localization and detection. In this paper, we propose a compositional model that leverages a new mid-level representation called compositional trajectories and a locally articulated spatiotemporal deformable parts model (LALSDPM) for fully action understanding. Our methods is advantageous in capturing the variable structure of dynamic human activity over a long range. First, the compositional trajectories capture long-ranging, frequently co-occurring groups of trajectories in space time and represent them in discriminative hierarchies, where human motion is largely separated from camera motion; second, LASTDPM learns a structured model with multi-layer deformable parts to capture multiple levels of articulated motion. We implement our methods and demonstrate state of the art performance on all three problems: action detection, localization, and recognition.