vix.ing · top · new · best · stats · spec

Affine-based Deformable Attention and Selective Fusion for Semi-dense Matching

2024/05/22 by Hongkai Chen, Zixin Luo, Chen, Hongkai +19 · 2 citations
Computer Science · Engineering · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face and Expression Recognition #Image Processing Techniques and Applications

paper · pdf · doi:10.48550/arxiv.2405.13874

openalex publication_date 2024/05/22 · openalex created_date 2024/05/26 · openalex updated_date 2026/07/28

Abstract

Identifying robust and accurate correspondences across images is a fundamental problem in computer vision that enables various downstream tasks. Recent semi-dense matching methods emphasize the effectiveness of fusing relevant cross-view information through Transformer. In this paper, we propose several improvements upon this paradigm. Firstly, we introduce affine-based local attention to model cross-view deformations. Secondly, we present selective fusion to merge local and global messages from cross attention. Apart from network structure, we also identify the importance of enforcing spatial smoothness in loss design, which has been omitted by previous works. Based on these augmentations, our network demonstrate strong matching capacity under different settings. The full version of our network achieves state-of-the-art performance among semi-dense matching methods at a similar cost to LoFTR, while the slim version reaches LoFTR baseline's performance with only 15% computation cost and 18% parameters.

Cited by

Related