2017/06/02 by Alexander Richard, Hilde Kuehne, Richard, Alexander +3
Computer Science · #Anomaly Detection Techniques and Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications #Video Analysis and Summarization
paper · pdf · doi:10.48550/arxiv.1706.00699
openalex publication_date 2017/06/02 · openalex created_date 2022/10/02 · openalex updated_date 2026/07/28
Action detection and temporal segmentation of actions in videos are topics of\nincreasing interest. While fully supervised systems have gained much attention\nlately, full annotation of each action within the video is costly and\nimpractical for large amounts of video data. Thus, weakly supervised action\ndetection and temporal segmentation methods are of great importance. While most\nworks in this area assume an ordered sequence of occurring actions to be given,\nour approach only uses a set of actions. Such action sets provide much less\nsupervision since neither action ordering nor the number of action occurrences\nare known. In exchange, they can be easily obtained, for instance, from\nmeta-tags, while ordered sequences still require human annotation. We introduce\na system that automatically learns to temporally segment and label actions in a\nvideo, where the only supervision that is used are action sets. An evaluation\non three datasets shows that our method still achieves good results although\nthe amount of supervision is significantly smaller than for other related\nmethods.\n