MatchAttention: Matching the Relative Positions for High-Resolution Cross-View Matching
2025/10/16 by Yan, Tingman, Liu, Tao, Yang, Xilian +2
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2510.14260
Abstract
Cross-view matching is fundamentally achieved through cross-attention mechanisms. However, matching of high-resolution images remains challenging due to the quadratic complexity and lack of explicit matching constraints in the existing cross-attention. This paper proposes an attention mechanism, MatchAttention, that dynamically matches relative positions. The relative position determines the attention sampling center of the key-value pairs given a query. Continuous and differentiable sliding-window attention sampling is achieved by the proposed BilinearSoftmax. The relative positions are iteratively updated through residual connections across layers by embedding them into the feature channels. Since the relative position is exactly the learning target for cross-view matching, an efficient hierarchical cross-view decoder, MatchDecoder, is designed with MatchAttention as its core component. To handle cross-view occlusions, gated cross-MatchAttention and a consistency-constrained loss are proposed. These two components collectively mitigate the impact of occlusions in both forward and backward passes, allowing the model to focus more on learning matching relationships. When applied to stereo matching, MatchStereo-B ranked 1st in average error on the public Middlebury benchmark and requires only 29ms for KITTI-resolution inference. MatchStereo-T can process 4K UHD images in 0.1 seconds using only 3GB of GPU memory. The proposed models also achieve state-of-the-art performance on KITTI 2012, KITTI 2015, ETH3D, and Spring flow datasets. The combination of high accuracy and low computational complexity makes real-time, high-resolution, and high-accuracy cross-view matching possible. Project page: https://github.com/TingmanYan/MatchAttention.
Citations
- StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
- Fast-FoundationStereo: Real-Time Zero-Shot Stereo Matching
- Lite Any Stereo: Efficient Zero-Shot Stereo Matching
- Stereo 3D Gaussian Splatting SLAM for Outdoor Urban Scenes
- What Makes Good Synthetic Training Data for Zero-Shot Stereo Matching?
- BANet: Bilateral Aggregation Network for Mobile Stereo Matching
- FoundationStereo: Zero-Shot Stereo Matching
- MonSter++: Unified Stereo Matching, Multi-view Stereo, and Real-time Stereo with Monodepth Priors
- ZeroStereo: Zero-shot Stereo Matching from Single Images
- CLEAR: Conv-Like Linearization Revs Pre-Trained Diffusion Transformers Up
- All-in-One: Transferring Vision Foundation Models into Stereo Matching
- Stereo Anywhere: Robust Zero-Shot Deep Stereo Matching Even Where Either Stereo or Mono Fail
- Binocular-Guided 3D Gaussian Splatting with View Consistency for Sparse View Synthesis
- Self-Evolving Depth-Supervised 3D Gaussian Splatting from Rendered Stereo Pairs
- IGEV++: Iterative Multi-range Geometry Encoding Volumes for Stereo Matching
- The Llama 3 Herd of Models
- LightStereo: Channel Boost Is All You Need for Efficient 2D Cost Aggregation
- Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
- Depth Anything V2
- SEA-RAFT: Simple, Efficient, Accurate RAFT for Optical Flow
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- MemFlow: Optical Flow Estimation and Prediction with Memory
- Selective-Stereo: Adaptive Frequency Information Selection for Stereo Matching
- DUSt3R: Geometric 3D Vision Made Easy
- TransNeXt: Robust Foveal Visual Perception for Vision Transformers
- Adaptive Multi-Modal Cross-Entropy Loss for Stereo Matching
- DINOv2: Learning Robust Visual Features without Supervision
- Micrograph segmentations for DDEVD
- Segment Anything
- Iterative Geometry Encoding Volume for Stereo Matching
- Spring: A High-Resolution High-Detail Dataset and Benchmark for Scene Flow, Optical Flow and Stereo
- A Practical Stereo Depth System for Smart Glasses
- CroCo v2: Improved Cross-view Completion Pre-training for Stereo Matching and Optical Flow
- Unifying Flow, Stereo and Depth Estimation
- MetaFormer Baselines for Vision
- SKFlow: Learning Optical Flow with Super Kernels
- Neighborhood Attention Transformer
- DIP: Deep Inverse Patchmatch for High-Resolution Optical Flow
- GraftNet: Towards Domain Generalized Stereo Matching with a Broad-Spectrum and Task-Oriented Feature
- CRAFT: Cross-Attentional Flow Transformer for Robust Optical Flow
- BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers
- FlowFormer: A Transformer Architecture for Optical Flow
- Practical Stereo Matching via Cascaded Recurrent Network with Adaptive Correlation
- Global Matching with Overlapping Attention for Optical Flow Estimation
- ITSA: An Information-Theoretic Approach to Automatic Shortcut Avoidance and Domain Generalization in Stereo Matching Networks
- GMFlow: Learning Optical Flow via Global Matching
- Swin Transformer V2: Scaling Up Capacity and Resolution
- RAFT-Stereo: Multilevel Recurrent Field Transforms for Stereo Matching
- Correlate-and-Excite: Real-Time Stereo Matching via Guided Cost Volume Excitation
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- CFNet: Cascade and Fused Cost Volume for Robust Stereo Matching
- Learning to Estimate Hidden Motions with Global Motion Aggregation
- Conditional Positional Encodings for Vision Transformers
- Deformable DETR: Deformable Transformers for End-to-End Object Detection
- HITNet: Hierarchical Iterative Tile Refinement Network for Real-time\n Stereo Matching
- Longformer: The Long-Document Transformer
- RAFT: Recurrent All-Pairs Field Transforms for Optical Flow
- RAFT: Recurrent All-Pairs Field Transforms for Optical Flow
- Virtual KITTI 2
- Domain-invariant Stereo Matching Networks
- Pyramid Stereo Matching Network
- PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume
- Attention Is All You Need
- End-to-End Learning of Geometry and Context for Deep Stereo Regression
- Gaussian Error Linear Units (GELUs)
- Effective Approaches to Attention-based Neural Machine Translation
- U-Net: Convolutional Networks for Biomedical Image Segmentation
- FlowNet: Learning Optical Flow with Convolutional Networks
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Neural Machine Translation by Jointly Learning to Align and Translate
- Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
- Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
Related