2017/07/28 by Pascal Mettes, Cees G. M. Snoek, Mettes, Pascal +1
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.1707.09145
openalex publication_date 2017/07/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We aim for zero-shot localization and classification of human actions in\nvideo. Where traditional approaches rely on global attribute or object\nclassification scores for their zero-shot knowledge transfer, our main\ncontribution is a spatial-aware object embedding. To arrive at spatial\nawareness, we build our embedding on top of freely available actor and object\ndetectors. Relevance of objects is determined in a word embedding space and\nfurther enforced with estimated spatial preferences. Besides local object\nawareness, we also embed global object awareness into our embedding to maximize\nactor and object interaction. Finally, we exploit the object positions and\nsizes in the spatial-aware embedding to demonstrate a new spatio-temporal\naction retrieval scenario with composite queries. Action localization and\nclassification experiments on four contemporary action video datasets support\nour proposal. Apart from state-of-the-art results in the zero-shot localization\nand classification settings, our spatial-aware embedding is even competitive\nwith recent supervised action localization alternatives.\n