SuperPoint: Self-Supervised Interest Point Detection and Description
2017/12/20 by Daniel DeTone, DeTone, Daniel, Tomasz Malisiewicz +3 · 1 voice · 252 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #cs.CV
paper · pdf · doi:10.48550/arxiv.1712.07629
Camera-ready version for CVPR 2018 Deep Learning for Visual SLAM Workshop (DL4VSLAM2018)
arxiv published 2017/12/20 · arxiv created 2018/04/19 · arxiv updated 2018/04/20
Abstract
This paper presents a self-supervised framework for training interest point detectors and descriptors suitable for a large number of multiple-view geometry problems in computer vision. As opposed to patch-based neural networks, our fully-convolutional model operates on full-sized images and jointly computes pixel-level interest point locations and associated descriptors in one forward pass. We introduce Homographic Adaptation, a multi-scale, multi-homography approach for boosting interest point detection repeatability and performing cross-domain adaptation (e.g., synthetic-to-real). Our model, when trained on the MS-COCO generic image dataset using Homographic Adaptation, is able to repeatedly detect a much richer set of interest points than the initial pre-adapted deep model and any other traditional corner detector. The final system gives rise to state-of-the-art homography estimation results on HPatches when compared to LIFT, SIFT and ORB.
Cited by
- Accuracy potential of visual localization exploiting high-end street-level imagery
- SC-Net: Robust Correspondence Learning via Spatial and Cross-Channel Context
- MGCA-Net: Multi-Graph Contextual Attention Network for Two-View Correspondence Learning
- Global-Aware Edge Prioritization for Pose Graph Initialization
- Long-tail Internet photo reconstruction
- Robust 6-DoF Object Pose Tracking with Built-In Recovery under Occlusions and Rapid Object Motions
- Match-Any-Events: Zero-Shot Motion-Robust Feature Matching Across Wide Baselines for Event Cameras
- Geometric Context Transformer for Streaming 3D Reconstruction
- Reloc-VGGT: Visual Re-localization with Geometry Grounded Transformer
- RANSAC Scoring Functions: Analysis and Reality Check
- Adaptive Covariance and Quaternion-Focused Hybrid Error-State EKF/UKF for Visual-Inertial Odometry
- Robust Scene Coordinate Regression via Geometrically-Consistent Global Descriptors
- NAP3D: NeRF Assisted 3D-3D Pose Alignment for Autonomous Vehicles
- Off The Grid: Detection of Primitives for Feed-Forward 3D Gaussian Splatting
- Cross-modal Fundus Image Registration under Large FoV Disparity
- Disturbance-Free Surgical Video Generation from Multi-Camera Shadowless Lamps for Open Surgery
- Geometry-Aware Sparse Depth Sampling for High-Fidelity RGB-D Depth Completion in Robotic Systems
- Trajectory Densification and Depth from Perspective-based Blur
- Selfi: Self Improving Reconstruction Engine via 3D Geometric Feature Alignment
- Automated Annotation of Shearographic Measurements Enabling Weakly Supervised Defect Detection
- FLIGHT: Fibonacci Lattice-based Inference for Geometric Heading in real-Time
- Cross-Stain Contrastive Learning for Paired Immunohistochemistry and Histopathology Slide Representation Learning
- What Is The Best 3D Scene Representation for Robotics? From Geometric to Foundation Models
- Emergent Outlier View Rejection in Visual Geometry Grounded Transformers
- MUT3R: Motion-aware Updating Transformer for Dynamic 3D Reconstruction
- Visual Sync: Multi-Camera Synchronization via Cross-View Object Motion
- MV-TAP: Tracking Any Point in Multi-View Videos
- EAG3R: Event-Augmented 3D Geometry Estimation for Dynamic and Extreme-Lighting Scenes
- ViGG: Robust RGB-D Point Cloud Registration using Visual-Geometric Mutual Guidance
- MARVO: Marine-Adaptive Radiance-aware Visual Odometry
- Emergent Extreme-View Geometry in 3D Foundation Models
- Gaussians on Fire: High-Frequency Reconstruction of Flames
- Semantic-Enhanced Feature Matching with Learnable Geometric Verification for Cross-Modal Neuron Registration
- Unlocking Zero-shot Potential of Semi-dense Image Matching via Gaussian Splatting
- C3Po: Cross-View Cross-Modality Correspondence by Pointmap Prediction
- SPIDER: Spatial Image CorresponDence Estimator for Robust Calibration
- NoPe-NeRF++: Local-to-Global Optimization of NeRF with No Pose Prior
- YOWO: You Only Walk Once to Jointly Map An Indoor Scene and Register Ceiling-mounted Cameras
- CLIDD: Cross-Layer Independent Deformable Description for Efficient and Discriminative Local Feature Representation
- Visible Structure Retrieval for Lightweight Image-Based Relocalisation
- DiffRegCD: Integrated Registration and Change Detection with Diffusion Features
- Wid3R: Wide Field-of-View 3D Reconstruction via Camera Model Conditioning
- U(PM)2:Unsupervised polygon matching with pre-trained models for challenging stereo images
- 4D3R: Motion-Aware Neural Reconstruction and Rendering of Dynamic Scenes from Monocular Videos
- DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular Videos
- GDROS: A Geometry-Guided Dense Registration Framework for Optical-SAR Images under Large Geometric Transformations
- Self-localization on a 3D map by fusing global and local features from a monocular camera
- SC-Match: Scale-Space Matching with Context Consistency for Side-Scan Sonar Mapping
- GeoComplete: Geometry-Aware Diffusion for Reference-Driven Image Completion
- Robust RPC Bundle Adjustment for Multi-Date Satellite Imagery with Season-Invariant Correspondences
- VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection
- RaCo: Ranking and Covariance for Practical Learned Keypoints
- Understanding and Optimizing Attention-Based Sparse Matching for Diverse Local Features
- XRefine: Attention-Guided Keypoint Match Refinement
- LightGlueStick: a Fast and Robust Glue for Joint Point-Line Matching
- InFlux: A Benchmark for Self-Calibration of Dynamic Intrinsics of Video Cameras
- FastJAM: a Fast Joint Alignment Model for Images
- LT-Exosense: A Vision-centric Multi-session Mapping System for Lifelong Safe Navigation of Exoskeletons
- Depth-Supervised Fusion Network for Seamless-Free Image Stitching
- PoseCrafter: Extreme Pose Estimation with Hybrid Video Synthesis
- Multi-Camera Worker Tracking in Logistics Warehouse Considering Wide-Angle Distortion
- MRASfM: Multi-Camera Reconstruction and Aggregation through Structure-from-Motion in Driving Scenes
- Leveraging AV1 motion vectors for Fast and Dense Feature Matching
- DeepDetect: Learning All-in-One Dense Keypoints
- CuSfM: CUDA-Accelerated Structure-from-Motion
- SaLon3R: Structure-aware Long-term Generalizable 3D Reconstruction from Unposed Images
- MACE: Mixture-of-Experts Accelerated Coordinate Encoding for Large-Scale Scene Localization and Rendering
- Through the Lens of Doubt: Robust and Efficient Uncertainty Estimation for Visual Place Recognition
- Trace Anything: Representing Any Video in 4D via Trajectory Fields
- Accelerated Feature Detectors for Visual SLAM: A Comparative Study of FPGA vs GPU
- ACE-G: Improving Generalization of Scene Coordinate Regression Through Query Pre-Training
- Guided Image Feature Matching using Feature Spatial Order
- DreamX-World 1.0: A General-Purpose Interactive World Model
- An Efficient Deep Template Matching and In-Plane Pose Estimation Method via Template-Aware Dynamic Convolution
- Dropping the D: RGB-D SLAM Without the Depth Sensor
- SegMASt3R: Geometry Grounded Segment Matching
- CLEAR-IR: Clarity-Enhanced Active Reconstruction of Infrared Imagery
- A Comparative Study of Vision Transformers and CNNs for Few-Shot Rigid Transformation and Fundamental Matrix Estimation
- OpenFLAME: Federated Visual Positioning System to Enable Large-Scale Augmented Reality Applications
- A Scene is Worth a Thousand Features: Feed-Forward Camera Localization from a Collection of Image Features
- HART: Human Aligned Reconstruction Transformer
- TTT3R: 3D Reconstruction as Test-Time Training
- GRS-SLAM3R: Real-Time Dense SLAM with Gated Recurrent State
- From Fields to Splats: A Cross-Domain Survey of Real-Time Neural Scene Representations
- RANSAC Scoring Done Right
- MASt3R-Fusion: Integrating Feed-Forward Visual Model with IMU, GNSS for High-Functionality SLAM
- GS-RoadPatching: Inpainting Gaussians via 3D Searching and Placing for Driving Scenes
- Aerial-Ground Image Feature Matching via 3D Gaussian Splatting-based Intermediate View Rendering
- Dark3R: Learning Structure from Motion in the Dark
- SAIL-Recon: Large SfM by Augmenting Scene Regression with Localization
- Camera Pose Refinement via 3D Gaussian Splatting
- KAMERA: Enhancing Aerial Surveys of Ice-associated Seals in Arctic Environments
- OrthoLoC: UAV 6-DoF Localization and Calibration Using Orthographic Geodata
- Tensor-Based Self-Calibration of Cameras via the TrifocalCalib Method
- SPFSplatV2: Efficient Self-Supervised Pose-Free 3D Gaussian Splatting from Sparse Views
- MS-GS: Multi-Appearance Sparse-View 3D Gaussian Splatting in the Wild
- Gaussian Alignment for Relative Camera Pose Estimation via Single-View Reconstruction
- OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
- A Geometrically Consistent Matching Framework for Side-Scan Sonar Mapping
- Efficient and Accurate Downfacing Visual Inertial Odometry
- Loc2: Interpretable Cross-View Localization via Depth-Lifted Local Feature Matching
- ObjectReact: Learning Object-Relative Control for Visual Navigation
- TANGO: Traversability-Aware Navigation with Local Metric Control for Topological Goals
- Gaussian Primitive Optimized Deformable Retinal Image Registration
- Good Deep Features to Track: Self-Supervised Feature Extraction and Tracking in Visual Odometry
- MOGS: Monocular Object-guided Gaussian Splatting in Large Scenes
- WinT3R: Window-Based Streaming Reconstruction with Camera Token Pool
- HAMSt3R: Human-Aware Multi-view Stereo 3D Reconstruction
- Aligned Anchor Groups Guided Line Segment Detector
- Radially Distorted Homographies, Revisited
- ActLoc: Learning to Localize on the Move via Active Viewpoint Selection
- Automated Feature Tracking for Real-Time Kinematic Analysis and Shape Estimation of Carbon Nanotube Growth
- Sensing, Social, and Motion Intelligence in Embodied Navigation: A Comprehensive Survey
- Remove360: Benchmarking Residuals After Object Removal in 3D Gaussian Splatting
- 3D Plant Root Skeleton Detection and Extraction
- EndoMatcher: Generalizable Endoscopic Image Matcher via Multi-Domain Pre-training for Robot-Assisted Surgery
- Refining Gaussian Splatting: A Volumetric Densification Approach
- COFFEE: A Shadow-Resilient Real-Time Pose Estimator for Unknown Tumbling Asteroids using Sparse Neural Networks
- SGAD: Semantic and Geometric-aware Descriptor for Local Feature Matching
- No Pose at All: Self-Supervised Pose-Free 3D Gaussian Splatting from Sparse Views
- GECO: Geometrically Consistent Embedding with Lightspeed Inference
- GeoMoE: Divide-and-Conquer Motion Field Modeling with Mixture-of-Experts for Two-View Geometry
- VMatcher: State-Space Semi-Dense Local Feature Matching
- Hi2-GSLoc: Dual-Hierarchical Gaussian-Specific Visual Relocalization for Remote Sensing
- Modality-Aware Feature Matching: A Comprehensive Review of Single- and Cross-Modality Techniques
- Estimating 2D Camera Motion with Hybrid Motion Basis
- Compress-Align-Detect: onboard change detection from unregistered images
- Impact of Underwater Image Enhancement on Feature Matching
- Reconstructing 4D Spatial Intelligence: A Survey
- PixelNav: Towards Model-based Vision-Only Navigation with Topological Graphs
- Cross Spatial Temporal Fusion Attention for Remote Sensing Object Detection via Image Feature Matching
- A 3D Cross-modal Keypoint Descriptor for MR-US Matching and Registration
- LONG3R: Long Sequence Streaming 3D Reconstruction
- LoopNet: A Multitasking Few-Shot Learning Approach for Loop Closure in Large Scale SLAM
- DSFormer: A Dual-Scale Cross-Learning Transformer for Visual Place Recognition
- C-DOG: Multi-View Multi-instance Feature Association Using Connected δ-Overlap Graphs
- CasP: Improving Semi-Dense Feature Matching Pipeline Leveraging Cascaded Correspondence Priors for Guidance
- VGGT-Long: Chunk it, Loop it, Align it -- Pushing VGGT's Limits on Kilometer-scale Long RGB Sequences
- A Single-step Accurate Fingerprint Registration Method Based on Local Feature Matching
- DINO-VO: A Feature-based Visual Odometry Leveraging a Visual Foundation Model
- CorrMoE: Mixture of Experts with De-stylization Learning for Cross-Scene and Cross-Domain Correspondence Pruning
- VISTA: Monocular Segmentation-Based Mapping for Appearance and View-Invariant Global Localization
- ScaleLSD: Scalable Deep Line Segment Detection Streamlined
- FPC-Net: Revisiting SuperPoint with Descriptor-Free Keypoint Detection via Feature Pyramids and Consistency-Based Implicit Matching
- Kaleidoscopic Background Attack: Disrupting Pose Estimation with Multi-Fold Radial Symmetry Textures
- Adaptive Framework for Ambient Intelligence in Rehabilitation Assistance
- PanMatch: Unleashing the Potential of Large Vision Models for Unified Matching Models
- Hardware-Aware Feature Extraction Quantisation for Real-Time Visual Odometry on FPGA Platforms
- Diff2I2P: Differentiable Image-to-Point Cloud Registration with Diffusion Prior
- 3DGSLSR:LargeScale Relocation for Autonomous Driving Based on 3D Gaussian Splatting
- RIPE: Reinforcement Learning on Unlabeled Image Pairs for Robust Keypoint Extraction
- U-ViLAR: Uncertainty-Aware Visual Localization for Autonomous Driving via Differentiable Association and Registration
- Grid-Reg: Detector-Free Gridized Feature Learning and Matching for Large-Scale SAR-Optical Image Registration
- Point3R: Streaming 3D Reconstruction with Explicit Spatial Pointer Memory
- Wildlife Target Re-Identification Using Self-supervised Learning in Non-Urban Settings
- LoD-Loc v2: Aerial Visual Localization over Low Level-of-Detail City Models using Explicit Silhouette Alignment
- Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space
- UFM: A Simple Path towards Unified Dense Correspondence with Flow
- Image Demoiréing Using Dual Camera Fusion on Mobile Phones
- Multimodal Fusion SLAM with Fourier Attention
- Co-VisiON: Co-Visibility ReasONing on Sparse Image Sets of Indoor Scenes
- LunarLoc: Segment-Based Global Localization on the Moon
- Semantic and Feature Guided Uncertainty Quantification of Visual Localization for Autonomous Vehicles
- Language-Grounded Hierarchical Planning and Execution with Multi-Robot 3D Scene Graphs
- Unsupervised Pelage Pattern Unwrapping for Animal Re-identification
- EmbodiedPlace: Learning Mixture-of-Features with Embodied Constraints for Visual Place Recognition
- SuperPoint-SLAM3: Augmenting ORB-SLAM3 with Deep Features, Adaptive NMS, and Learning-Based Loop Closure
- Test3R: Learning to Reconstruct 3D at Test Time
- Unsupervised Deformable Image Registration with Structural Nonparametric Smoothing
- Multimodal Spatial Language Maps for Robot Navigation and Manipulation
- RealKeyMorph: Keypoints in Real-world Coordinates for Resolution-agnostic Image Registration
- Hierarchical Image Matching for UAV Absolute Visual Localization via Semantic and Structural Constraints
- Splat and Replace: 3D Reconstruction with Repetitive Elements
- OXSeg: Multidimensional attention UNet-based lip segmentation using semi-supervised lip contours
- O-MaMa: Learning Object Mask Matching between Egocentric and Exocentric Views
- CryoFastAR: Fast Cryo-EM Ab Initio Reconstruction Made Easy
- Active Illumination Control in Low-Light Environments using NightHawk
- Deep Learning Reforms Image Matching: A Survey and Outlook
- SupeRANSAC: One RANSAC to Rule Them All
- Accelerating SfM-based Pose Estimation with Dominating Set
- SplArt: Articulation Estimation and Part-Level Reconstruction with 3D Gaussian Splatting
- Seeing in the Dark: Benchmarking Egocentric 3D Vision with the Oxford Day-and-Night Dataset
- Dense Match Summarization for Faster Two-view Estimation
- VolTex: Food Volume Estimation using Text-Guided Segmentation and Neural Surface Reconstruction
- Implicit Deformable Medical Image Registration with Learnable Kernels
- SAB3R: Semantic-Augmented Backbone in 3D Reconstruction
- SteerPose: Simultaneous Extrinsic Camera Calibration and Matching from Articulation
- Rig3R: Rig-Aware Conditioning for Learned 3D Reconstruction
- Flying Co-Stereo: Enabling Long-Range Aerial Dense Mapping via Collaborative Stereo Vision of Dynamic-Baseline
- Black-box Adversarial Attacks on CNN-based SLAM Algorithms
- LiftFeat: 3D Geometry-Aware Local Feature Matching
- Rooms from Motion: Un-posed Indoor 3D Object Detection as Localization and Mapping
- TimePoint: Accelerated Time Series Alignment via Self-Supervised Keypoint and Descriptor Learning
- Unsupervised training of keypoint-agnostic descriptors for flexible retinal image registration
- Unsupervised Deep Learning-based Keypoint Localization Estimating Descriptor Matching Performance
- TwinTrack: Bridging Vision and Contact Physics for Real-Time Tracking of Unknown Dynamic Objects
- UAVPairs: A Challenging Benchmark for Match Pair Retrieval of Large-scale UAV Images
- Collaborative Learning for Unsupervised Multimodal Remote Sensing Image Registration: Integrating Self-Supervision and MIM-Guided Diffusion-Based Image Translation
- Visual Loop Closure Detection Through Deep Graph Consensus
- Focus What Matters: Matchability-Based Reweighting for Local Feature Matching
- Why Not Replace? Sustaining Long-Term Visual Localization via Handcrafted-Learned Feature Collaboration on CPU
- A Coarse to Fine 3D LiDAR Localization with Deep Local Features for Long Term Robot Navigation in Large Environments
- To Glue or Not to Glue? Classical vs Learned Image Matching for Mobile Mapping Cameras to Textured Semantic 3D Building Models
- Decoupled Geometric Parameterization and its Application in Deep Homography Estimation
- Incorporating Visual Correspondence into Diffusion Model for Virtual Try-On
- GMatch: A Lightweight, Geometry-Constrained Keypoint Matcher for Zero-Shot 6DoF Pose Estimation in Robotic Grasp Tasks
- 3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering
- Learning Cross-Spectral Point Features with Task-Oriented Training
- Large-Scale Gaussian Splatting SLAM
- AeroMap3D: Anchoring Monocular UAV 6-DoF Localization to Visual-Geometric-Semantic Map Priors
- DS-SAC: Density Search for Sample Consensus
- RDD: Robust Feature Detector and Descriptor using Deformable Transformer
- A Simple Detector with Frame Dynamics is a Strong Tracker
- Auto-regressive transformation for image alignment
- DiffusionSfM: Predicting Structure and Motion via Ray Origin and Endpoint Diffusion
- Enhancing Lidar Point Cloud Sampling via Colorization and Super-Resolution of Lidar Imagery
- OBD-Finder: Explainable Coarse-to-Fine Text-Centric Oracle Bone Duplicates Discovery
- Argus: Metric Panoramic 3D Reconstruction for Indoor Scenes
- T-Graph: Enhancing Sparse-view Camera Pose Estimation by Pairwise Translation Graph
- Are Minimal Radial Distortion Solvers Really Necessary for Relative Pose Estimation?
- OVGGT: O(1) Constant-Cost Streaming Visual Geometry Transformer
- Efficient Dense Matching for Enhanced Gaussian Splatting Using AV1 Motion Vectors
- Unconstrained Large-scale 3D Reconstruction and Rendering across Altitudes
- GSFeatLoc: Visual Localization Using Feature Correspondence on 3D Gaussian Splatting
- Large-scale visual SLAM for in-the-wild videos
- Marginalized Bundle Adjustment: Multi-View Camera Pose from Monocular Depth Estimates
- MP-SfM: Monocular Surface Priors for Robust Structure-from-Motion
- Demonstrating DVS: Dynamic Virtual-Real Simulation Platform for Mobile Robotic Tasks
- ImLoc: Revisiting Visual Localization with Image-based Representation
- SGFormer: Structure-Guided Transformer for Robust Local Feature Matching
- Dynamic Camera Poses and Where to Find Them
- EdgePoint2: Compact Descriptors for Superior Efficiency and Accuracy
- A Guide to Structureless Visual Localization
- Differential 6-DOF Pose Estimation with Provable First-Order Immunity to Camera Calibration Errors
- LoRetta: A Foundation Model and Extensive Dataset for Global-Scale Remote Sensing Dense Image Matching
- Label-Free Target-Domain Adaptation for Unconstrained Event-Image Feature Matching via Dual-Stage Distillation
- PRaDA: Projective Radial Distortion Averaging
- An Accelerated Camera 3DMA Framework for Efficient Urban GNSS Multipath Estimation
- SmallGS: Gaussian Splatting-based Camera Pose Estimation for Small-Baseline Videos
- SELC: Self-Supervised Efficient Local Correspondence Learning for Low Quality Images
- Multimodal Perception for Goal-oriented Navigation: A Survey
- Pose Optimization for Autonomous Driving Datasets using Neural Rendering Models
- Exploring Generalizable Pre-training for Real-world Change Detection via Geometric Estimation
- Mono3R: Exploiting Monocular Cues for Geometric 3D Reconstruction
- AerialMegaDepth: Learning Aerial-Ground Reconstruction and View Synthesis
- Are Pretrained Image Matchers Good Enough for SAR-Optical Satellite Registration?
- Knowledge Distillation for Underwater Feature Extraction and Matching via GAN-synthesized Images
- Hardware, Algorithms, and Applications of the Neuromorphic Vision Sensor: a Review
- GraspClutter6D: A Large-scale Real-world Dataset for Robust Perception and Grasping in Cluttered Scenes
- POMATO: Marrying Pointmap Matching with Temporal Motion for Dynamic 3D Reconstruction
- To Match or Not to Match: Revisiting Image Matching for Reliable Visual Place Recognition
- Learning Affine Correspondences by Integrating Geometric Constraints
Discussions
Related