Distinctive Image Features from Scale-Invariant Keypoints
2004/06/02 by David Lowe, David G. Lowe · 257 citations
Computer Science · Engineering · #Advanced Image and Video Retrieval Techniques #Image and Object Detection Techniques #Robotics and Sensor-Based Localization
paper · doi:10.1023/b:visi.0000029664.99615.94
openalex publication_date 2004/06/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/31
Cited by
- Gaussians on Fire: High-Frequency Reconstruction of Flames
- Multimodal remote sensing change detection: An image matching perspective
- Hybrid SIFT-SNN for Efficient Anomaly Detection of Traffic Flow-Control Infrastructure
- Unlocking Zero-shot Potential of Semi-dense Image Matching via Gaussian Splatting
- Selective Disk Bispectrum: A Complete and Rotation Invariant Image Descriptor
- Real-Time Object Tracking with On-Device Deep Learning for Adaptive Beamforming in Dynamic Acoustic Environments
- C3Po: Cross-View Cross-Modality Correspondence by Pointmap Prediction
- FeRA: Frequency-Energy Constrained Routing for Effective Diffusion Adaptation Fine-Tuning
- SPIDER: Spatial Image CorresponDence Estimator for Robust Calibration
- iGaussian: Real-Time Camera Pose Estimation via Feed-Forward 3D Gaussian Splatting Inversion
- Find the Leak, Fix the Split: Cluster-Based Method to Prevent Leakage in Video-Derived Datasets
- Uni-Hand: Universal Hand Motion Forecasting in Egocentric Views
- CLIDD: Cross-Layer Independent Deformable Description for Efficient and Discriminative Local Feature Representation
- DiffRegCD: Integrated Registration and Change Detection with Diffusion Features
- Wid3R: Wide Field-of-View 3D Reconstruction via Camera Model Conditioning
- LeCoT: revisiting network architecture for two-view correspondence pruning
- Global Multiple Extraction Network for Low-Resolution Facial Expression Recognition
- Improving Multi-View Reconstruction via Texture-Guided Gaussian-Mesh Joint Optimization
- MID: A Self-supervised Multimodal Iterative Denoising Framework
- VisionCAD: An Integration-Free Radiology Copilot Framework
- Scaling Image Geo-Localization to Continent Level
- Robust RPC Bundle Adjustment for Multi-Date Satellite Imagery with Season-Invariant Correspondences
- VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection
- RaCo: Ranking and Covariance for Practical Learned Keypoints
- Understanding and Optimizing Attention-Based Sparse Matching for Diverse Local Features
- XRefine: Attention-Guided Keypoint Match Refinement
- High-throughput Verticillium wilt detection in cotton: A comparative study of faster R-CNN and YOLOv11
- Hybrid Vision Servoing with Depp Alignment and GRU-Based Occlusion Recovery
- GroundLoc: Efficient Large-Scale Outdoor LiDAR-Only Localization
- Reanimating the past: From historical collections of the placenta and uterus to modern imaging, machine learning, and multiscale modeling
- LightGlueStick: a Fast and Robust Glue for Joint Point-Line Matching
- PlanarTrack: A high-quality and challenging benchmark for large-scale planar object tracking
- Dynamically Detect and Fix Hardness for Efficient Approximate Nearest Neighbor Search
- Depth-Supervised Fusion Network for Seamless-Free Image Stitching
- Freehand 3D Ultrasound Imaging: Sim-in-the-Loop Probe Pose Optimization via Visual Servoing
- UREM: A High-performance Unified and Resilient Enhancement Method for Multi- and High-Dimensional Indexes
- PoseCrafter: Extreme Pose Estimation with Hybrid Video Synthesis
- Ninja Codes: Neurally Generated Fiducial Markers for Stealthy 6-DoF Tracking
- A Renaissance of Explicit Motion Information Mining from Transformers for Action Recognition
- Joint Multi-Condition Representation Modelling via Matrix Factorisation for Visual Place Recognition
- Leveraging AV1 motion vectors for Fast and Dense Feature Matching
- DeepDetect: Learning All-in-One Dense Keypoints
- CuSfM: CUDA-Accelerated Structure-from-Motion
- C4D: 4D Made from 3D through Dual Correspondences
- Exploring Image Representation with Decoupled Classical Visual Descriptors
- Leveraging Cycle-Consistent Anchor Points for Self-Supervised RGB-D Registration
- Scaling Vision Transformers for Functional MRI with Flat Maps
- Through the Lens of Doubt: Robust and Efficient Uncertainty Estimation for Visual Place Recognition
- Accelerated Feature Detectors for Visual SLAM: A Comparative Study of FPGA vs GPU
- MultiFoodhat: A potential new paradigm for intelligent food quality inspection
- Guided Image Feature Matching using Feature Spatial Order
- An Efficient Deep Template Matching and In-Plane Pose Estimation Method via Template-Aware Dynamic Convolution
- LTGS: Long-Term Gaussian Scene Chronology From Sparse View Updates
- The Orbitoscope, a six-axis macro-imaging robot for photogrammetric 3D-digitization of insects and other small specimens
- A Comparative Study of Vision Transformers and CNNs for Few-Shot Rigid Transformation and Fundamental Matrix Estimation
- OpenFLAME: Federated Visual Positioning System to Enable Large-Scale Augmented Reality Applications
- Panorama: Fast-Track Nearest Neighbors
- Reward driven discovery of the optimal microstructure representations with invariant variational autoencoders
- Benchmarking Egocentric Visual-Inertial SLAM at City Scale
- Hy-Facial: Hybrid Feature Extraction by Dimensionality Reduction Methods for Enhanced Facial Expression Classification
- SAGE: Spatial-visual Adaptive Graph Exploration for Visual Place Recognition
- TTT3R: 3D Reconstruction as Test-Time Training
- Robust Visual Localization in Compute-Constrained Environments by Salient Edge Rendering and Weighted Hamming Similarity
- Similarity-Aware Selective State-Space Modeling for Semantic Correspondence
- Adaptive Canonicalization with Application to Invariant Anisotropic Geometric Networks
- RANSAC Scoring Done Right
- OracleGS: Grounding Generative Priors for Sparse-View Gaussian Splatting
- ARD-REFSM: Enhancing Reflection Symmetry Detection with Asymmetric Denoising and Rotation Equivariance
- Counting Grid Aggregation for Event Retrieval and Recognition
- Scalable and Sustainable Dry Microfabrication Enabled by High-Precision and Wafer-Scale Transfer Lithography of Commercial Photoresists
- Dark3R: Learning Structure from Motion in the Dark
- Camera Pose Refinement via 3D Gaussian Splatting
- Identifying Adaptive Footprints in the Presence of Demographic Uncertainty
- Effective scaling registration approach by imposing the emphasis on the scale factor
- KAMERA: Enhancing Aerial Surveys of Ice-associated Seals in Arctic Environments
- Hierarchical Neural Semantic Representation for 3D Semantic Correspondence
- OrthoLoC: UAV 6-DoF Localization and Calibration Using Orthographic Geodata
- Automatic Intermodal Loading Unit Identification using Computer Vision: A Scoping Review
- SPFSplatV2: Efficient Self-Supervised Pose-Free 3D Gaussian Splatting from Sparse Views
- HotSpotter - Patterned Species Instance Recognition
- Aberrant cortex contractions impact mammalian oocyte quality
- Geometric Image Synchronization with Deep Watermarking
- PRISM: Product Retrieval In Shopping Carts using Hybrid Matching
- Scale and Rotation Estimation of Similarity-Transformed Images via Cross-Correlation Maximization Based on Auxiliary Function Method
- SWA-PF: Semantic-Weighted Adaptive Particle Filter for Memory-Efficient 4-DoF UAV Localization in GNSS-Denied Environments
- Gaussian Alignment for Relative Camera Pose Estimation via Single-View Reconstruction
- MFAF: An EVA02-Based Multi-scale Frequency Attention Fusion Method for Cross-View Geo-Localization
- OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
- The Gap of Semantic Parsing: A Survey on Automatic Math Word Problem Solvers
- Towards the Distributed Large-scale k-NN Graph Construction by Graph Merge
- Calib3R: A 3D Foundation Model for Multi-Camera to Robot Calibration and 3D Metric-Scaled Scene Reconstruction
- Deep Learning for Free-Hand Sketch: A Survey
- Good Deep Features to Track: Self-Supervised Feature Extraction and Tracking in Visual Odometry
- Understanding the Limitations of CNN-based Absolute Camera Pose\n Regression
- Robotic Tactile Perception of Object Properties: A Review
- EDFFDNet: Towards Accurate and Efficient Unsupervised Multi-Grid Image Registration
- Australian Supermarket Object Set (ASOS): A Benchmark Dataset of Physical Objects and 3D Models for Robotics and Computer Vision
- Block-Sparse Global Attention for Efficient Multi-View Geometry Transformers
- Learning Multi-Stage Tasks with One Demonstration via Self-Replay
- TemporalFlowViz: Parameter-Aware Visual Analytics for Interpreting Scramjet Combustion Evolution
- Comparative Evaluation of Traditional and Deep Learning Feature Matching Algorithms using Chandrayaan-2 Lunar Data
- Towards Open World Detection: A Survey
- Stitching the Story: Creating Panoramic Incident Summaries from Body-Worn Footage
- Global-to-Local or Local-to-Global? Enhancing Image Retrieval with Efficient Local Search and Effective Global Re-ranking
- On sources to variabilities of simple cells in the primary visual cortex: A principled theory for the interaction between geometric image transformations and receptive field responses
- MOSAIC: Multi-Subject Personalized Generation via Correspondence-Aware Alignment and Disentanglement
- Articulated Object Estimation in the Wild
- Multi-Focused Video Group Activities Hashing
- Encoder-Only Image Registration
- GelSLAM: A Real-time, High-Fidelity, and Robust 3D Tactile SLAM System
- Teaching Robots to Do Object Assembly using Multi-modal 3D Vision
- Radially Distorted Homographies, Revisited
- Real-time Collaboration Between Mixed Reality Users in Geo-referenced Virtual Environment
- Looking Beyond the Obvious: A Survey on Abstract Concept Recognition for Video Understanding
- MSMVD: Exploiting Multi-scale Image Features via Multi-scale BEV Features for Multi-view Pedestrian Detection
- Seam360GS: Seamless 360° Gaussian Splatting from Real-World Omnidirectional Images
- Automated Feature Tracking for Real-Time Kinematic Analysis and Shape Estimation of Carbon Nanotube Growth
- Spatial Policy: Guiding Visuomotor Robotic Manipulation with Spatial-Aware Modeling and Reasoning
- CellINR: Implicitly Overcoming Photo-induced Artifacts in 4D Live Fluorescence Microscopy
- LIFT: Learned Invariant Feature Transform
- Minimally Supervised Feature Selection for Classification (Master's Thesis, University Politehnica of Bucharest)
- HEp-2 Cell Image Classification with Deep Convolutional Neural Networks
- Reliable Multi-view 3D Reconstruction for `Just-in-time' Edge Environments
- Sensing, Social, and Motion Intelligence in Embodied Navigation: A Comprehensive Survey
- A Generic Framework for Assessing the Performance Bounds of Image\n Feature Detectors
- Large-Scale Image Retrieval with Attentive Deep Local Features
- Opening the Black Box of 3D Reconstruction Error Analysis with VECTOR
- Local Scale Equivariance with Latent Deep Equilibrium Canonicalizer
- Viewpoint Invariant Object Detector
- Matching objects across the textured-smooth continuum
- Deep learning-driven investigation of nanoplastic impacts on soil protist behavior in soil chips
- A smile I could recognise in a thousand: Automatic identification of identity from dental radiography
- Graph-based Thermal-Inertial SLAM with Probabilistic Neural Networks
- Learning to Zoom: a Saliency-Based Sampling Layer for Neural Networks
- Data Shift of Object Detection in Autonomous Driving
- Feature Learning by Multidimensional Scaling and its Applications in Object Recognition
- Neural View Synthesis and Matching for Semi-Supervised Few-Shot Learning of 3D Pose
- Neural Probabilistic System for Text Recognition
- A comparative study of plant phenotyping workflows based on three-dimensional reconstruction from multi-view images
- Axis-level Symmetry Detection with Group-Equivariant Representation
- Multi-View Large-Scale Bundle Adjustment Method for High-Resolution Satellite Images
- A Sub-Pixel Multimodal Optical Remote Sensing Images Matching Method
- Deep Sparse Subspace Clustering
- Compact Approximation for Polynomial of Covariance Feature
- Convolutional Hough Matching Networks for Robust and Efficient Visual Correspondence
- LiPo-LCD: Combining Lines and Points for Appearance-based Loop Closure Detection
- Topological Structure Description for Artcode Detection Using the Shape of Orientation Histogram
- Image Retrieval with Fisher Vectors of Binary Features
- LATCH: Learned Arrangements of Three Patch Codes
- Fast, Accurate Thin-Structure Obstacle Detection for Autonomous Mobile Robots
- ImageNet MPEG-7 Visual Descriptors - Technical Report
- Salient Object Detection: A Distinctive Feature Integration Model
- PatchGame: Learning to Signal Mid-level Patches in Referential Games
- Harnessing Input-Adaptive Inference for Efficient VLN
- Learning to Hallucinate Face Images via Component Generation and Enhancement
- Sparse and redundant signal representations for x-ray computed tomography
- A Performance Evaluation of Local Features for Image Based 3D Reconstruction
- Improving Place Recognition Using Dynamic Object Detection
- Mem4D: Decoupling Static and Dynamic Memory for Dynamic Scene Reconstruction
- Semi-supervised Multiscale Matching for SAR-Optical Image
- Dynamicity and Durability in Scalable Visual Instance Search
- Are visual dictionaries generalizable?
- Deep Learning for Single-View Instance Recognition
- Efficient Multimedia Similarity Measurement Using Similar Elements
- Dynamic Concept Composition for Zero-Example Event Detection
- Robust Image Stitching with Optimal Plane
- EndoMatcher: Generalizable Endoscopic Image Matcher via Multi-Domain Pre-training for Robot-Assisted Surgery
- TSMS-SAM2: Multi-scale Temporal Sampling Augmentation and Memory-Splitting Pruning for Promptable Video Object Segmentation and Tracking in Surgical Scenarios
- Separable Four Points Fundamental Matrix
- Deep Learning-based Scalable Image-to-3D Facade Parser for Generating Thermal 3D Building Models
- PIS3R: Very Large Parallax Image Stitching via Deep 3D Reconstruction
- Asymmetric activation of retinal ON and OFF pathways by AOSLO raster-scanned visual stimuli
- DOMR: Establishing Cross-View Segmentation via Dense Object Matching
- Infrared face recognition: a comprehensive review of methodologies and\n databases
- Content-Based Bird Retrieval using Shape context, Color moments and Bag of Features
- Unsupervised Metric Relocalization Using Transform Consistency Loss
- COFFEE: A Shadow-Resilient Real-Time Pose Estimator for Unknown Tumbling Asteroids using Sparse Neural Networks
- A Review of Visual Odometry Methods and Its Applications for Autonomous Driving
- Learned versus Hand-Designed Feature Representations for 3d Agglomeration
- Contextual Action Recognition with R*CNN
- Learning Spread-out Local Feature Descriptors
- SGAD: Semantic and Geometric-aware Descriptor for Local Feature Matching
- Evolutionary Paradigms in Histopathology Serial Sections technology
- CVD-SfM: A Cross-View Deep Front-end Structure-from-Motion System for Sparse Localization in Multi-Altitude Scenes
- Saliency based Semi-supervised Learning for Orbiting Satellite Tracking
- GeoMoE: Divide-and-Conquer Motion Field Modeling with Mixture-of-Experts for Two-View Geometry
- Weakly Supervised Virus Capsid Detection with Image-Level Annotations in Electron Microscopy Images
- Learn by Observation: Imitation Learning for Drone Patrolling from Videos of A Human Navigator
- Towards Robust Semantic Correspondence: A Benchmark and Insights
- 3D Reconstruction via Incremental Structure From Motion
- A Data-driven Approach for Human Pose Tracking Based on Spatio-temporal Pictorial Structure
- Working hard to know your neighbor's margins: Local descriptor learning loss
- VMatcher: State-Space Semi-Dense Local Feature Matching
- MatchBench: An Evaluation of Feature Matchers
- Modality-Aware Feature Matching: A Comprehensive Review of Single- and Cross-Modality Techniques
- Estimating 2D Camera Motion with Hybrid Motion Basis
- Efficient Spatial-Temporal Modeling for Real-Time Video Analysis: A Unified Framework for Action Recognition and Object Tracking
- Moiré Zero: An Efficient and High-Performance Neural Architecture for Moiré Removal
- PanoSplatt3R: Leveraging Perspective Pretraining for Generalized Unposed Wide-Baseline Panorama Reconstruction
- Detection Method Based on Automatic Visual Shape Clustering for Pin-Missing Defect in Transmission Lines
- Image Recognition Using Scale Recurrent Neural Networks
- Deep Learning Representation using Autoencoder for 3D Shape Retrieval
- Impact of Underwater Image Enhancement on Feature Matching
- A Reconstruction Error Formulation for Semi-Supervised Multi-task and Multi-view Learning
- ST-DAI: Single-shot 2.5D Spatial Transcriptomics with Intra-Sample Domain Adaptive Imputation for Cost-efficient 3D Reconstruction
- Visual Odometry Revisited: What Should Be Learnt?
- Vision-based system identification and 3D keypoint discovery using dynamics constraints
- ProNet: Learning to Propose Object-specific Boxes for Cascaded Neural Networks
- Reconstructing Sinus Anatomy from Endoscopic Video -- Towards a Radiation-free Approach for Quantitative Longitudinal Assessment
- RAID: A Relation-Augmented Image Descriptor
- Aggregating Deep Convolutional Features for Image Retrieval
- A Sparse Coding Interpretation of Neural Networks and Theoretical Implications
- SK-Net: Deep Learning on Point Cloud via End-to-end Discovery of Spatial Keypoints
- Computer-Aided Colorectal Tumor Classification in NBI Endoscopy Using CNN Features
- 4D Cardiac Ultrasound Standard Plane Location by Spatial-Temporal Correlation
- Robust SfM with Little Image Overlap
- ADAGIO: Fast Data-aware Near-Isometric Linear Embeddings
- 6-DoF Object Pose from Semantic Keypoints
- GIFT: Learning Transformation-Invariant Dense Visual Descriptors via Group CNNs
- FRAM: Frobenius-Regularized Assignment Matching with Mixed-Precision Computing
- Cross Spatial Temporal Fusion Attention for Remote Sensing Object Detection via Image Feature Matching
- Target Driven Instance Detection
- TROVE Feature Detection for Online Pose Recovery by Binocular Cameras
- Superpixel-based Two-view Deterministic Fitting for Multiple-structure Data
- Pose Graph Optimization for Unsupervised Monocular Visual Odometry
- Kernel Selection using Multiple Kernel Learning and Domain Adaptation in Reproducing Kernel Hilbert Space, for Face Recognition under Surveillance Scenario
- Improving the HardNet Descriptor
- Graph based Nearest Neighbor Search: Promises and Failures
- End-to-end learning of keypoint detection and matching for relative pose estimation
- LONG3R: Long Sequence Streaming 3D Reconstruction
- VLASE: Vehicle Localization by Aggregating Semantic Edges
- Towards an Automatic System for Extracting Planar Orientations from Software Generated Point Clouds
- Window-Object Relationship Guided Representation Learning for Generic Object Detections
- cvpaper.challenge in 2016: Futuristic Computer Vision through 1,600 Papers Survey
- DCTM: Discrete-Continuous Transformation Matching for Semantic Flow
- Robust Place Recognition using an Imaging Lidar
- Using depth information and colour space variations for improving outdoor robustness for instance segmentation of cabbage
- Megalithic statue (moai) production on Rapa Nui (Easter Island, Chile)
- Joint Spatial and Layer Attention for Convolutional Networks
- Joint Deformable Registration of Large EM Image Volumes: A Matrix Solver\n Approach
- Sparseness helps: Sparsity Augmented Collaborative Representation for Classification
- Retrieval and Localization with Observation Constraints
- Understanding Deep Image Representations by Inverting Them
- Local Unsupervised Learning for Image Analysis
- Visual-inertial navigation, mapping and localization: A scalable real-time causal approach
- Scene Invariant Crowd Segmentation and Counting Using Scale-Normalized\n Histogram of Moving Gradients (HoMG)
- Learning Multi-Scale Deep Features for High-Resolution Satellite Image Classification
- Generative Adversarial Data Programming
- Image‐based 3D Modelling: A Review
- <scp>3D LiDAR SLAM</scp>: A survey
- A local fingerprinting approach for audio copy detection
- Adaptive color-corrected multicolor space enhancement network for underwater image enhancement
- Microassembly of multi-material and 3D integration enabled by programmable and universal high-precision micro-transfer printing
- Optimising Image Feature Extraction and Selection: A Comprehensive Review With Spark Case Studies
- Evaluating Soccer Player: from Live Camera to Deep Reinforcement Learning
- A Benchmark Comparison of Visual Place Recognition Techniques for Resource-Constrained Embedded Platforms
- An Intelligent Hybrid Model for Identity Document Classification