2021/11/10 by Zhengyi Liu, Yacheng Tan, Qian He +1 · 360 citations
Computer Science · Engineering · #Advanced Neural Network Applications #Artificial intelligence #Computer science #Computer vision #Edge detection #Electrical engineering #Engineering #Image (mathematics) #Image processing #RGB color model #Transformer #Video Surveillance and Tracking Methods #Visual Attention and Saliency Detection #Voltage #cs.CV
paper · pdf · doi:10.1109/tcsvt.2021.3127149
published in IEEE Transactions on Circuits and Systems for Video Technology 32(7), 4486-4497 (Institute of Electrical and Electronics Engineers) · Online published in TCSVT
openalex publication_date 2021/11/10 · arxiv created 2022/04/12 · arxiv updated 2022/04/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Convolutional neural networks (CNNs) are good at extracting contexture features within certain receptive fields, while transformers can model the global long-range dependency features. By absorbing the advantage of transformer and the merit of CNN, Swin Transformer shows strong feature representation ability. Based on it, we propose a cross-modality fusion model SwinNet for RGB-D and RGB-T salient object detection. It is driven by Swin Transformer to extract the hierarchical features, boosted by attention mechanism to bridge the gap between two modalities, and guided by edge information to sharp the contour of salient object. To be specific, two-stream Swin Transformer encoder first extracts multi-modality features, and then spatial alignment and channel re-calibration module is presented to optimize intra-level cross-modality features. To clarify the fuzzy boundary, edge-guided decoder achieves inter-level cross-modality fusion under the guidance of edge features. The proposed model outperforms the state-of-the-art models on RGB-D and RGB-T datasets, showing that it provides more insight into the cross-modality complementarity task.