vix.ing · top · new · best · stats · spec

An end-to-end generative framework for video segmentation and\n recognition

2015/09/07 by Hilde Kuehne, Kuehne, Hilde, Jüergen Gall +3 · 3 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Human Pose and Action Recognition #Video Analysis and Summarization #Video Surveillance and Tracking Methods

paper · pdf · doi:10.48550/arxiv.1509.01947

openalex publication_date 2015/09/07 · openalex created_date 2022/10/01 · openalex updated_date 2026/07/28

Abstract

We describe an end-to-end generative approach for the segmentation and\nrecognition of human activities. In this approach, a visual representation\nbased on reduced Fisher Vectors is combined with a structured temporal model\nfor recognition. We show that the statistical properties of Fisher Vectors make\nthem an especially suitable front-end for generative models such as Gaussian\nmixtures. The system is evaluated for both the recognition of complex\nactivities as well as their parsing into action units. Using a variety of video\ndatasets ranging from human cooking activities to animal behaviors, our\nexperiments demonstrate that the resulting architecture outperforms\nstate-of-the-art approaches for larger datasets, i.e. when sufficient amount of\ndata is available for training structured generative models.\n

Cited by

Related