2017/08/17 by Jae Shin Yoon, François Rameau, Yoon, Jae Shin +9 · 3 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Visual Attention and Saliency Detection
paper · pdf · doi:10.48550/arxiv.1708.05137
openalex publication_date 2017/08/17 · openalex created_date 2022/10/01 · openalex updated_date 2026/07/28
We propose a novel video object segmentation algorithm based on pixel-level\nmatching using Convolutional Neural Networks (CNN). Our network aims to\ndistinguish the target area from the background on the basis of the pixel-level\nsimilarity between two object units. The proposed network represents a target\nobject using features from different depth layers in order to take advantage of\nboth the spatial details and the category-level semantic information.\nFurthermore, we propose a feature compression technique that drastically\nreduces the memory requirements while maintaining the capability of feature\nrepresentation. Two-stage training (pre-training and fine-tuning) allows our\nnetwork to handle any target object regardless of its category (even if the\nobject's type does not belong to the pre-training data) or of variations in its\nappearance through a video sequence. Experiments on large datasets demonstrate\nthe effectiveness of our model - against related methods - in terms of\naccuracy, speed, and stability. Finally, we introduce the transferability of\nour network to different domains, such as the infrared data domain.\n