2021/01/31 by Josip Šarić, Sacha Vražić, Siniša Šegvić · 6 citations
Computer Science · Engineering · Mathematics · #Anomaly Detection Techniques and Applications #Artificial intelligence #Computer science #Computer vision #Engineering #Feature (linguistics) #Generative Adversarial Networks and Image Synthesis #Joint (building) #Mathematics #Motion (physics) #Pattern recognition (psychology) #Regression #Semantic feature #Statistics #Video Analysis and Summarization #cs.CV
paper · pdf · doi:10.1109/tnnls.2021.3136624
published in IEEE Transactions on Neural Networks and Learning Systems 34(9), 6443-6455 (Institute of Electrical and Electronics Engineers) · 13 pages, 10 figures
arxiv created 2021/12/16 · openalex publication_date 2021/12/28 · arxiv updated 2022/01/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Dense semantic forecasting anticipates future events in the video by inferring pixel-level semantics of an unobserved future image. We present a novel approach that is applicable to various single-frame architectures and tasks. Our approach consists of two modules. The feature-to-motion (F2M) module forecasts a dense deformation field that warps past features into their future positions. The feature-to-feature (F2F) module regresses the future features directly and is, therefore, able to account for emergent scenery. The compound F2MF model decouples the effects of motion from the effects of novelty in a task-agnostic manner. We aim to apply F2MF forecasting to the most subsampled and the most abstract representation of the desired single-frame model. Our design takes advantage of deformable convolutions and spatial correlation coefficients across neighboring time instants. We perform experiments on three dense prediction tasks: semantic segmentation, instance-level segmentation, and panoptic segmentation. The results reveal state-of-the-art forecasting accuracy across three dense prediction tasks.