Distinctive Image Features from Scale-Invariant Keypoints
2004/06/02 by David Lowe, David G. Lowe · 984 citations
Computer Science · Engineering · #Advanced Image and Video Retrieval Techniques #Image and Object Detection Techniques #Robotics and Sensor-Based Localization
paper · doi:10.1023/b:visi.0000029664.99615.94
openalex publication_date 2004/06/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/31
Cited by
- Gaussians on Fire: High-Frequency Reconstruction of Flames
- Multimodal remote sensing change detection: An image matching perspective
- Hybrid SIFT-SNN for Efficient Anomaly Detection of Traffic Flow-Control Infrastructure
- Unlocking Zero-shot Potential of Semi-dense Image Matching via Gaussian Splatting
- Selective Disk Bispectrum: A Complete and Rotation Invariant Image Descriptor
- Real-Time Object Tracking with On-Device Deep Learning for Adaptive Beamforming in Dynamic Acoustic Environments
- C3Po: Cross-View Cross-Modality Correspondence by Pointmap Prediction
- FeRA: Frequency-Energy Constrained Routing for Effective Diffusion Adaptation Fine-Tuning
- SPIDER: Spatial Image CorresponDence Estimator for Robust Calibration
- iGaussian: Real-Time Camera Pose Estimation via Feed-Forward 3D Gaussian Splatting Inversion
- Find the Leak, Fix the Split: Cluster-Based Method to Prevent Leakage in Video-Derived Datasets
- Uni-Hand: Universal Hand Motion Forecasting in Egocentric Views
- CLIDD: Cross-Layer Independent Deformable Description for Efficient and Discriminative Local Feature Representation
- DiffRegCD: Integrated Registration and Change Detection with Diffusion Features
- Wid3R: Wide Field-of-View 3D Reconstruction via Camera Model Conditioning
- LeCoT: revisiting network architecture for two-view correspondence pruning
- Global Multiple Extraction Network for Low-Resolution Facial Expression Recognition
- Improving Multi-View Reconstruction via Texture-Guided Gaussian-Mesh Joint Optimization
- MID: A Self-supervised Multimodal Iterative Denoising Framework
- VisionCAD: An Integration-Free Radiology Copilot Framework
- DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular Videos
- Scaling Image Geo-Localization to Continent Level
- Robust RPC Bundle Adjustment for Multi-Date Satellite Imagery with Season-Invariant Correspondences
- VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection
- RaCo: Ranking and Covariance for Practical Learned Keypoints
- Understanding and Optimizing Attention-Based Sparse Matching for Diverse Local Features
- XRefine: Attention-Guided Keypoint Match Refinement
- High-throughput Verticillium wilt detection in cotton: A comparative study of faster R-CNN and YOLOv11
- Hybrid Vision Servoing with Depp Alignment and GRU-Based Occlusion Recovery
- GroundLoc: Efficient Large-Scale Outdoor LiDAR-Only Localization
- Reanimating the past: From historical collections of the placenta and uterus to modern imaging, machine learning, and multiscale modeling
- LightGlueStick: a Fast and Robust Glue for Joint Point-Line Matching
- PlanarTrack: A high-quality and challenging benchmark for large-scale planar object tracking
- Dynamically Detect and Fix Hardness for Efficient Approximate Nearest Neighbor Search
- Depth-Supervised Fusion Network for Seamless-Free Image Stitching
- Freehand 3D Ultrasound Imaging: Sim-in-the-Loop Probe Pose Optimization via Visual Servoing
- UREM: A High-performance Unified and Resilient Enhancement Method for Multi- and High-Dimensional Indexes
- PoseCrafter: Extreme Pose Estimation with Hybrid Video Synthesis
- Ninja Codes: Neurally Generated Fiducial Markers for Stealthy 6-DoF Tracking
- A Renaissance of Explicit Motion Information Mining from Transformers for Action Recognition
- Joint Multi-Condition Representation Modelling via Matrix Factorisation for Visual Place Recognition
- Leveraging AV1 motion vectors for Fast and Dense Feature Matching
- DeepDetect: Learning All-in-One Dense Keypoints
- CuSfM: CUDA-Accelerated Structure-from-Motion
- C4D: 4D Made from 3D through Dual Correspondences
- Exploring Image Representation with Decoupled Classical Visual Descriptors
- Leveraging Cycle-Consistent Anchor Points for Self-Supervised RGB-D Registration
- Scaling Vision Transformers for Functional MRI with Flat Maps
- Through the Lens of Doubt: Robust and Efficient Uncertainty Estimation for Visual Place Recognition
- Accelerated Feature Detectors for Visual SLAM: A Comparative Study of FPGA vs GPU
- MultiFoodhat: A potential new paradigm for intelligent food quality inspection
- Guided Image Feature Matching using Feature Spatial Order
- An Efficient Deep Template Matching and In-Plane Pose Estimation Method via Template-Aware Dynamic Convolution
- LTGS: Long-Term Gaussian Scene Chronology From Sparse View Updates
- The Orbitoscope, a six-axis macro-imaging robot for photogrammetric 3D-digitization of insects and other small specimens
- A Comparative Study of Vision Transformers and CNNs for Few-Shot Rigid Transformation and Fundamental Matrix Estimation
- OpenFLAME: Federated Visual Positioning System to Enable Large-Scale Augmented Reality Applications
- Panorama: Fast-Track Nearest Neighbors
- Reward driven discovery of the optimal microstructure representations with invariant variational autoencoders
- Benchmarking Egocentric Visual-Inertial SLAM at City Scale
- Hy-Facial: Hybrid Feature Extraction by Dimensionality Reduction Methods for Enhanced Facial Expression Classification
- SAGE: Spatial-visual Adaptive Graph Exploration for Efficient Visual Place Recognition
- TTT3R: 3D Reconstruction as Test-Time Training
- Robust Visual Localization in Compute-Constrained Environments by Salient Edge Rendering and Weighted Hamming Similarity
- Similarity-Aware Selective State-Space Modeling for Semantic Correspondence
- Adaptive Canonicalization with Application to Invariant Anisotropic Geometric Networks
- RANSAC Scoring Done Right
- OracleGS: Grounding Generative Priors for Sparse-View Gaussian Splatting
- ARD-REFSM: Enhancing Reflection Symmetry Detection with Asymmetric Denoising and Rotation Equivariance
- Counting Grid Aggregation for Event Retrieval and Recognition
- Scalable and Sustainable Dry Microfabrication Enabled by High-Precision and Wafer-Scale Transfer Lithography of Commercial Photoresists
- Dark3R: Learning Structure from Motion in the Dark
- Camera Pose Refinement via 3D Gaussian Splatting
- Identifying Adaptive Footprints in the Presence of Demographic Uncertainty
- Effective scaling registration approach by imposing the emphasis on the scale factor
- KAMERA: Enhancing Aerial Surveys of Ice-associated Seals in Arctic Environments
- Hierarchical Neural Semantic Representation for 3D Semantic Correspondence
- OrthoLoC: UAV 6-DoF Localization and Calibration Using Orthographic Geodata
- Automatic Intermodal Loading Unit Identification using Computer Vision: A Scoping Review
- SPFSplatV2: Efficient Self-Supervised Pose-Free 3D Gaussian Splatting from Sparse Views
- HotSpotter - Patterned Species Instance Recognition
- Aberrant cortex contractions impact mammalian oocyte quality
- Geometric Image Synchronization with Deep Watermarking
- PRISM: Product Retrieval In Shopping Carts using Hybrid Matching
- Scale and Rotation Estimation of Similarity-Transformed Images via Cross-Correlation Maximization Based on Auxiliary Function Method
- SWA-PF: Semantic-Weighted Adaptive Particle Filter for Memory-Efficient 4-DoF UAV Localization in GNSS-Denied Environments
- Gaussian Alignment for Relative Camera Pose Estimation via Single-View Reconstruction
- MFAF: An EVA02-Based Multi-scale Frequency Attention Fusion Method for Cross-View Geo-Localization
- OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
- The Gap of Semantic Parsing: A Survey on Automatic Math Word Problem Solvers
- Towards the Distributed Large-scale k-NN Graph Construction by Graph Merge
- Calib3R: A 3D Foundation Model for Multi-Camera to Robot Calibration and 3D Metric-Scaled Scene Reconstruction
- Deep Learning for Free-Hand Sketch: A Survey
- Good Deep Features to Track: Self-Supervised Feature Extraction and Tracking in Visual Odometry
- Understanding the Limitations of CNN-based Absolute Camera Pose Regression
- Robotic Tactile Perception of Object Properties: A Review
- EDFFDNet: Towards Accurate and Efficient Unsupervised Multi-Grid Image Registration
- Australian Supermarket Object Set (ASOS): A Benchmark Dataset of Physical Objects and 3D Models for Robotics and Computer Vision
- Block-Sparse Global Attention for Efficient Multi-View Geometry Transformers
- Learning Multi-Stage Tasks with One Demonstration via Self-Replay
- TemporalFlowViz: Parameter-Aware Visual Analytics for Interpreting Scramjet Combustion Evolution
- Comparative Evaluation of Traditional and Deep Learning Feature Matching Algorithms using Chandrayaan-2 Lunar Data
- Towards Open World Detection: A Survey
- Stitching the Story: Creating Panoramic Incident Summaries from Body-Worn Footage
- Global-to-Local or Local-to-Global? Enhancing Image Retrieval with Efficient Local Search and Effective Global Re-ranking
- On sources to variabilities of simple cells in the primary visual cortex: A principled theory for the interaction between geometric image transformations and receptive field responses
- MOSAIC: Multi-Subject Personalized Generation via Correspondence-Aware Alignment and Disentanglement
- Articulated Object Estimation in the Wild
- Multi-Focused Video Group Activities Hashing
- Encoder-Only Image Registration
- GelSLAM: A Real-time, High-Fidelity, and Robust 3D Tactile SLAM System
- Teaching Robots to Do Object Assembly using Multi-modal 3D Vision
- Radially Distorted Homographies, Revisited
- Real-time Collaboration Between Mixed Reality Users in Geo-referenced Virtual Environment
- Looking Beyond the Obvious: A Survey on Abstract Concept Recognition for Video Understanding
- MSMVD: Exploiting Multi-scale Image Features via Multi-scale BEV Features for Multi-view Pedestrian Detection
- Seam360GS: Seamless 360° Gaussian Splatting from Real-World Omnidirectional Images
- Automated Feature Tracking for Real-Time Kinematic Analysis and Shape Estimation of Carbon Nanotube Growth
- Spatial Policy: Guiding Visuomotor Robotic Manipulation with Spatial-Aware Modeling and Reasoning
- CellINR: Implicitly Overcoming Photo-induced Artifacts in 4D Live Fluorescence Microscopy
- LIFT: Learned Invariant Feature Transform
- Minimally Supervised Feature Selection for Classification (Master's Thesis, University Politehnica of Bucharest)
- HEp-2 Cell Image Classification with Deep Convolutional Neural Networks
- Reliable Multi-view 3D Reconstruction for `Just-in-time' Edge Environments
- Sensing, Social, and Motion Intelligence in Embodied Navigation: A Comprehensive Survey
- A Generic Framework for Assessing the Performance Bounds of Image Feature Detectors
- Large-Scale Image Retrieval with Attentive Deep Local Features
- Opening the Black Box of 3D Reconstruction Error Analysis with VECTOR
- Local Scale Equivariance with Latent Deep Equilibrium Canonicalizer
- Viewpoint Invariant Object Detector
- Matching objects across the textured-smooth continuum
- Deep learning-driven investigation of nanoplastic impacts on soil protist behavior in soil chips
- A smile I could recognise in a thousand: Automatic identification of identity from dental radiography
- Graph-based Thermal-Inertial SLAM with Probabilistic Neural Networks
- Learning to Zoom: a Saliency-Based Sampling Layer for Neural Networks
- Data Shift of Object Detection in Autonomous Driving
- Feature Learning by Multidimensional Scaling and its Applications in Object Recognition
- Neural View Synthesis and Matching for Semi-Supervised Few-Shot Learning of 3D Pose
- Neural Probabilistic System for Text Recognition
- A comparative study of plant phenotyping workflows based on three-dimensional reconstruction from multi-view images
- Axis-level Symmetry Detection with Group-Equivariant Representation
- Multi-View Large-Scale Bundle Adjustment Method for High-Resolution Satellite Images
- A Sub-Pixel Multimodal Optical Remote Sensing Images Matching Method
- Deep Sparse Subspace Clustering
- Compact Approximation for Polynomial of Covariance Feature
- Convolutional Hough Matching Networks for Robust and Efficient Visual Correspondence
- LiPo-LCD: Combining Lines and Points for Appearance-based Loop Closure Detection
- Topological Structure Description for Artcode Detection Using the Shape of Orientation Histogram
- Image Retrieval with Fisher Vectors of Binary Features
- LATCH: Learned Arrangements of Three Patch Codes
- Fast, Accurate Thin-Structure Obstacle Detection for Autonomous Mobile Robots
- ImageNet MPEG-7 Visual Descriptors - Technical Report
- Salient Object Detection: A Distinctive Feature Integration Model
- PatchGame: Learning to Signal Mid-level Patches in Referential Games
- Harnessing Input-Adaptive Inference for Efficient VLN
- Learning to Hallucinate Face Images via Component Generation and Enhancement
- Sparse and redundant signal representations for x-ray computed tomography
- A Performance Evaluation of Local Features for Image Based 3D Reconstruction
- Improving Place Recognition Using Dynamic Object Detection
- Mem4D: Decoupling Static and Dynamic Memory for Dynamic Scene Reconstruction
- Semi-supervised Multiscale Matching for SAR-Optical Image
- Dynamicity and Durability in Scalable Visual Instance Search
- Active Self Calibration of a Multi Sensor System
- Are visual dictionaries generalizable?
- Deep Learning for Single-View Instance Recognition
- Efficient Multimedia Similarity Measurement Using Similar Elements
- Dynamic Concept Composition for Zero-Example Event Detection
- Robust Image Stitching with Optimal Plane
- EndoMatcher: Generalizable Endoscopic Image Matcher via Multi-Domain Pre-training for Robot-Assisted Surgery
- TSMS-SAM2: Multi-scale Temporal Sampling Augmentation and Memory-Splitting Pruning for Promptable Video Object Segmentation and Tracking in Surgical Scenarios
- Separable Four Points Fundamental Matrix
- Combined Image Data Augmentations diminish the benefits of Adaptive Label Smoothing
- Deep Learning-based Scalable Image-to-3D Facade Parser for Generating Thermal 3D Building Models
- PIS3R: Very Large Parallax Image Stitching via Deep 3D Reconstruction
- Asymmetric activation of retinal ON and OFF pathways by AOSLO raster-scanned visual stimuli
- DOMR: Establishing Cross-View Segmentation via Dense Object Matching
- Deep Convolutional Neural Network for 6-DOF Image Localization
- Infrared face recognition: a comprehensive review of methodologies and databases
- Content-Based Bird Retrieval using Shape context, Color moments and Bag of Features
- Unsupervised Metric Relocalization Using Transform Consistency Loss
- COFFEE: A Shadow-Resilient Real-Time Pose Estimator for Unknown Tumbling Asteroids using Sparse Neural Networks
- Unsupervised Learning of Visual Representations using Videos
- A Review of Visual Odometry Methods and Its Applications for Autonomous Driving
- Learned versus Hand-Designed Feature Representations for 3d Agglomeration
- Contextual Action Recognition with R*CNN
- Learning Spread-out Local Feature Descriptors
- SGAD: Semantic and Geometric-aware Descriptor for Local Feature Matching
- Evolutionary Paradigms in Histopathology Serial Sections technology
- CVD-SfM: A Cross-View Deep Front-end Structure-from-Motion System for Sparse Localization in Multi-Altitude Scenes
- Saliency based Semi-supervised Learning for Orbiting Satellite Tracking
- Robust Photogeometric Localization over Time for Map-Centric Loop Closure
- GeoMoE: Divide-and-Conquer Motion Field Modeling with Mixture-of-Experts for Two-View Geometry
- Weakly Supervised Virus Capsid Detection with Image-Level Annotations in Electron Microscopy Images
- Learn by Observation: Imitation Learning for Drone Patrolling from Videos of A Human Navigator
- Towards Robust Semantic Correspondence: A Benchmark and Insights
- 3D Reconstruction via Incremental Structure From Motion
- A Data-driven Approach for Human Pose Tracking Based on Spatio-temporal Pictorial Structure
- Working hard to know your neighbor's margins: Local descriptor learning loss
- VMatcher: State-Space Semi-Dense Local Feature Matching
- MatchBench: An Evaluation of Feature Matchers
- Modality-Aware Feature Matching: A Comprehensive Review of Single- and Cross-Modality Techniques
- Estimating 2D Camera Motion with Hybrid Motion Basis
- Efficient Spatial-Temporal Modeling for Real-Time Video Analysis: A Unified Framework for Action Recognition and Object Tracking
- Moiré Zero: An Efficient and High-Performance Neural Architecture for Moiré Removal
- Planar Object Tracking in the Wild: A Benchmark
- PanoSplatt3R: Leveraging Perspective Pretraining for Generalized Unposed Wide-Baseline Panorama Reconstruction
- Detection Method Based on Automatic Visual Shape Clustering for Pin-Missing Defect in Transmission Lines
- Image Recognition Using Scale Recurrent Neural Networks
- Deep Learning Representation using Autoencoder for 3D Shape Retrieval
- Impact of Underwater Image Enhancement on Feature Matching
- A Reconstruction Error Formulation for Semi-Supervised Multi-task and Multi-view Learning
- ST-DAI: Single-shot 2.5D Spatial Transcriptomics with Intra-Sample Domain Adaptive Imputation for Cost-efficient 3D Reconstruction
- Visual Odometry Revisited: What Should Be Learnt?
- Vision-based system identification and 3D keypoint discovery using dynamics constraints
- ProNet: Learning to Propose Object-specific Boxes for Cascaded Neural Networks
- Reconstructing Sinus Anatomy from Endoscopic Video -- Towards a Radiation-free Approach for Quantitative Longitudinal Assessment
- RAID: A Relation-Augmented Image Descriptor
- Aggregating Deep Convolutional Features for Image Retrieval
- Cascade Network with Guided Loss and Hybrid Attention for Two-view Geometry
- A Sparse Coding Interpretation of Neural Networks and Theoretical Implications
- SK-Net: Deep Learning on Point Cloud via End-to-end Discovery of Spatial Keypoints
- Computer-Aided Colorectal Tumor Classification in NBI Endoscopy Using CNN Features
- 4D Cardiac Ultrasound Standard Plane Location by Spatial-Temporal Correlation
- Robust SfM with Little Image Overlap
- Geometric Processing for Image-based 3D Object Modeling
- A Reinforcement Learning Approach for Sequential Spatial Transformer Networks
- Explaining First Impressions: Modeling, Recognizing, and Explaining Apparent Personality from Videos
- Fast k-means based on KNN Graph
- ADAGIO: Fast Data-aware Near-Isometric Linear Embeddings
- 6-DoF Object Pose from Semantic Keypoints
- GIFT: Learning Transformation-Invariant Dense Visual Descriptors via Group CNNs
- FRAM: Frobenius-Regularized Assignment Matching with Mixed-Precision Computing
- Cross Spatial Temporal Fusion Attention for Remote Sensing Object Detection via Image Feature Matching
- Target Driven Instance Detection
- TROVE Feature Detection for Online Pose Recovery by Binocular Cameras
- Superpixel-based Two-view Deterministic Fitting for Multiple-structure Data
- Pose Graph Optimization for Unsupervised Monocular Visual Odometry
- Kernel Selection using Multiple Kernel Learning and Domain Adaptation in Reproducing Kernel Hilbert Space, for Face Recognition under Surveillance Scenario
- Improving the HardNet Descriptor
- Graph based Nearest Neighbor Search: Promises and Failures
- End-to-end learning of keypoint detection and matching for relative pose estimation
- LONG3R: Long Sequence Streaming 3D Reconstruction
- VLASE: Vehicle Localization by Aggregating Semantic Edges
- Towards an Automatic System for Extracting Planar Orientations from Software Generated Point Clouds
- Window-Object Relationship Guided Representation Learning for Generic Object Detections
- LoopNet: A Multitasking Few-Shot Learning Approach for Loop Closure in Large Scale SLAM
- cvpaper.challenge in 2016: Futuristic Computer Vision through 1,600 Papers Survey
- DCTM: Discrete-Continuous Transformation Matching for Semantic Flow
- Robust Place Recognition using an Imaging Lidar
- Using depth information and colour space variations for improving outdoor robustness for instance segmentation of cabbage
- Megalithic statue (moai) production on Rapa Nui (Easter Island, Chile)
- Joint Spatial and Layer Attention for Convolutional Networks
- Joint Deformable Registration of Large EM Image Volumes: A Matrix Solver Approach
- Sparseness helps: Sparsity Augmented Collaborative Representation for Classification
- Retrieval and Localization with Observation Constraints
- Understanding Deep Image Representations by Inverting Them
- Local Unsupervised Learning for Image Analysis
- Visual-inertial navigation, mapping and localization: A scalable real-time causal approach
- Scene Invariant Crowd Segmentation and Counting Using Scale-Normalized Histogram of Moving Gradients (HoMG)
- Learning Multi-Scale Deep Features for High-Resolution Satellite Image Classification
- An Evaluation of DUSt3R/MASt3R/VGGT 3D Reconstruction on Photogrammetric Aerial Blocks
- Generative Adversarial Data Programming
- Image‐based 3D Modelling: A Review
- 3D LiDAR SLAM: A survey
- A local fingerprinting approach for audio copy detection
- Adaptive color-corrected multicolor space enhancement network for underwater image enhancement
- Microassembly of multi-material and 3D integration enabled by programmable and universal high-precision micro-transfer printing
- Optimising Image Feature Extraction and Selection: A Comprehensive Review With Spark Case Studies
- Evaluating Soccer Player: from Live Camera to Deep Reinforcement Learning
- A Benchmark Comparison of Visual Place Recognition Techniques for Resource-Constrained Embedded Platforms
- An Intelligent Hybrid Model for Identity Document Classification
- Feature Detection for Hand Hygiene Stages
- What Is the Best Practice for CNNs Applied to Visual Instance Retrieval?
- Recent Advances in 3D Object and Hand Pose Estimation
- Image-to-GPS Verification Through A Bottom-Up Pattern Matching Network
- Surpassing Human-Level Face Verification Performance on LFW with GaussianFace
- Iterated Support Vector Machines for Distance Metric Learning
- SENNS: Sparse Extraction Neural NetworkS for Feature Extraction
- Geo-distinctive Visual Element Matching for Location Estimation of Images
- GPU Accelerated Cascade Hashing Image Matching for Large Scale 3D Reconstruction
- Chromatic Aberration Recovery on Arbitrary Images
- Temporally Robust Global Motion Compensation by Keypoint-based Congealing
- Person Re-Identification via Recurrent Feature Aggregation
- ACNe: Attentive Context Normalization for Robust Permutation-Equivariant Learning
- MKL-RT: Multiple Kernel Learning for Ratio-trace Problems via Convex Optimization
- Perception-based energy functions in seam-cutting
- X-ray Scattering Image Classification Using Deep Learning
- Foundations of Vector Retrieval
- Optical Navigation in Unstructured Dynamic Railroad Environments
- Where is my Phone ? Personal Object Retrieval from Egocentric Images
- Developing efficient transfer learning strategies for robust scene recognition in mobile robotics using pre-trained convolutional neural networks
- Discriminative Learning of Similarity and Group Equivariant Representations
- ScreenAvoider: Protecting Computer Screens from Ubiquitous Cameras
- Accurate Localization in Dense Urban Area Using Google Street View Image
- Spectrum Translation for Cross-Spectral Ocular Matching
- Frankenstein: Learning Deep Face Representations using Small Data
- Three-dimensional models of natural environments and the mapping of navigational information
- A Robust Indoor Scene Recognition Method based on Sparse Representation
- Backtracking Regression Forests for Accurate Camera Relocalization
- Approximate Similarity Search for Online Multimedia Services on Distributed CPU-GPU Platforms
- Self-Taught Support Vector Machine
- Image-based Vehicle Classification System
- Image classification based on support vector machine and the fusion of complementary features
- Object Detection Using Keygraphs
- Simultaneous Feature Aggregating and Hashing for Large-scale Image Search
- On Relative Pose Recovery for Multi-Camera Systems
- To go deep or wide in learning?
- DXSLAM: A Robust and Efficient Visual SLAM System with Deep Features
- CasP: Improving Semi-Dense Feature Matching Pipeline Leveraging Cascaded Correspondence Priors for Guidance
- DANIEL: A Fast and Robust Consensus Maximization Method for Point Cloud Registration with High Outlier Ratios
- ‘Structure-from-Motion’ photogrammetry: A low-cost, effective tool for geoscience applications
- Features extraction for image identification using computer vision
- Carl-Hauser -- Open Source Image Matching Algorithms Benchmarking Framework
- FishNet: A Camera Localizer using Deep Recurrent Networks
- Screen Gleaning: A Screen Reading TEMPEST Attack on Mobile Devices Exploiting an Electromagnetic Side Channel
- Bilinear Random Projections for Locality-Sensitive Binary Codes
- VolumeDeform: Real-time Volumetric Non-rigid Reconstruction
- Hybrid Indexes to Expedite Spatial-Visual Search
- A Discriminative Learned CNN Embedding for Remote Sensing Image Scene Classification
- CorrMoE: Mixture of Experts with De-stylization Learning for Cross-Scene and Cross-Domain Correspondence Pruning
- Detecting Human Interventions on the Landscape: KAZE Features, Poisson Point Processes, and a Construction Dataset
- SpatialTrackerV2: 3D Point Tracking Made Easy
- Image Quality Assessment Guided Deep Neural Networks Training
- VISTA: Monocular Segmentation-Based Mapping for Appearance and View-Invariant Global Localization
- Stochastic Dykstra Algorithms for Metric Learning on Positive Semi-Definite Cone
- Episodic Training for Domain Generalization
- Learning Edge-Preserved Image Stitching from Large-Baseline Deep Homography
- Sequence to Sequence Learning for Optical Character Recognition
- Embedding based on function approximation for large scale image search
- A Machine-Synesthetic Approach To DDoS Network Attack Detection
- Exploiting Contextual Information with Deep Neural Networks
- Computer-Assisted Analysis of Biomedical Images
- Tracking using Numerous Anchor points
- ViewSynth: Learning Local Features from Depth using View Synthesis
- Stacked Quantizers for Compositional Vector Compression
- Optimized hierarchical block matching for fast and accurate image registration
- Invariance of visual operations at the level of receptive fields
- Superimage
- Are State-of-the-art Visual Place Recognition Techniques any Good for Aerial Robotics?
- SConE: Siamese Constellation Embedding Descriptor for Image Matching
- Geometry-aware Similarity Learning on SPD Manifolds for Visual Recognition
- Histograms of Oriented Gradients for Landmine Detection in Ground-Penetrating Radar Data
- HomographyAD: Deep Anomaly Detection Using Self Homography Learning
- Douglas-Quaid -- Open Source Image Matching Library
- Deep Attentional Structured Representation Learning for Visual Recognition
- Dynamic Scale Inference by Entropy Minimization
- DR Loss: Improving Object Detection by Distributional Ranking
- A Keygraph Classification Framework for Real-Time Object Detection
- An Empirical Study of Non-Rigid Surface Feature Matching of Human from 3D Video
- Nonnegative Restricted Boltzmann Machines for Parts-based Representations Discovery and Predictive Model Stabilization
- Pose from Shape: Deep Pose Estimation for Arbitrary 3D Objects
- Privacy-preserving Learning via Deep Net Pruning
- Learning to Augment Expressions for Few-shot Fine-grained Facial Expression Recognition
- Fine-Grained Texture Identification for Reliable Product Traceability
- SPL-MLL: Selecting Predictable Landmarks for Multi-Label Learning
- Reliability Validation of Learning Enabled Vehicle Tracking
- Harvesting, Detecting, and Characterizing Liver Lesions from Large-scale Multi-phase CT Data via Deep Dynamic Texture Learning
- Vanishing Point Guided Natural Image Stitching
- Circulant temporal encoding for video retrieval and temporal alignment
- P3: Toward Privacy-Preserving Photo Sharing
- Multimodal Icon Annotation For Mobile Applications
- Multi-Person Pose Estimation with Enhanced Feature Aggregation and Selection
- Multi-Class Detection and Segmentation of Objects in Depth
- Artificial Intelligence for Pediatric Ophthalmology
- STN-Homography: estimate homography parameters directly
- Deep Reinforcement Learning in Computer Vision: A Comprehensive Survey
- Adversarial Robustness Guarantees for Classification with Gaussian Processes
- Shape Primitive Histogram: A Novel Low-Level Face Representation for Face Recognition
- Event Specific Multimodal Pattern Mining with Image-Caption Pairs
- Factorized Topic Models
- An Automatic Image Content Retrieval Method for better Mobile Device\n Display User Experiences
- A Machine Learning Approach to Recovery of Scene Geometry from Images
- FPC-Net: Revisiting SuperPoint with Descriptor-Free Keypoint Detection via Feature Pyramids and Consistency-Based Implicit Matching
- Aggregated Residual Transformations for Deep Neural Networks
- Improving Gamma-ray Source Search with Image Processing
- Who and Where: People and Location Co-Clustering
- Compact 3D Map-Based Monocular Localization Using Semantic Edge Alignment
- Convolutional Recurrent Residual U-Net Embedded with Attention Mechanism and Focal Tversky Loss Function for Cancerous Nuclei Detection
- Accurate Motion Estimation through Random Sample Aggregated Consensus
- Modern Physiognomy: An Investigation on Predicting Personality Traits and Intelligence from the Human Face
- Joint Max Margin and Semantic Features for Continuous Event Detection in Complex Scenes
- A Jointly Learned Deep Architecture for Facial Attribute Analysis and Face Detection in the Wild
- Weight-Sharing Neural Architecture Search: A Battle to Shrink the Optimization Gap
- Learning to Predict Repeatability of Interest Points
- Hardware-Aware Feature Extraction Quantisation for Real-Time Visual Odometry on FPGA Platforms
- 6-PACK: Category-level 6D Pose Tracker with Anchor-Based Keypoints
- MESS: Fast and Private Semantic Search on Multi-Graph HNSW
- Attribute-Graph: A Graph based approach to Image Ranking
- Nested Graph Words for Object Recognition
- Learnable Pooling Regions for Image Classification
- Learning to detect and localize many objects from few examples
- Fine-Grained Product Class Recognition for Assisted Shopping
- Detecting change in graffiti using a hybrid framework
- On the Costs and Benefits of Learned Indexing for Dynamic High-Dimensional Data: Extended Version
- Learning to Reconstruct and Segment 3D Objects
- Building A Large Concept Bank for Representing Events in Video
- Recent Advance in Content-based Image Retrieval: A Literature Survey
- Generative Panoramic Image Stitching
- Deep Epitomic Convolutional Neural Networks
- CorrelationFlow: A Training-Free Geometric Approach for LiDAR Scene Flow Estimation
- Salient Object Detection Combining a Self-attention Module and a Feature Pyramid Network
- RIPE: Reinforcement Learning on Unlabeled Image Pairs for Robust Keypoint Extraction
- SimPatch: A Nearest Neighbor Similarity Match between Image Patches
- An Experimental Study of Deep Convolutional Features For Iris Recognition
- Object Detection in 20 Years: A Survey
- Single-Shot Clothing Category Recognition in Free-Configurations with Application to Autonomous Clothes Sorting
- Cross-Modal and Multimodal Data Analysis Based on Functional Mapping of Spectral Descriptors and Manifold Regularization
- On the Reliability of Profile Matching Across Large Online Social Networks
- LiveChess2FEN: a Framework for Classifying Chess Pieces based on CNNs
- U-ViLAR: Uncertainty-Aware Visual Localization for Autonomous Driving via Differentiable Association and Registration
- Multi-Expert Learning Framework with the State Space Model for Optical and SAR Image Registration
- Auto-JacoBin: Auto-encoder Jacobian Binary Hashing
- DASC: Robust Dense Descriptor for Multi-modal and Multi-spectral Correspondence Estimation
- Outdoor Monocular SLAM with Global Scale-Consistent 3D Gaussian Pointmaps
- The CUDA LATCH Binary Descriptor: Because Sometimes Faster Means Better
- MGSfM: Multi-Camera Geometry Driven Global Structure-from-Motion
- Efficient Neighbourhood Consensus Networks via Submanifold Sparse Convolutions
- A Brief Survey and an Application of Semantic Image Segmentation for Autonomous Driving
- LoopDB: A Loop Closure Dataset for Large Scale Simultaneous Localization and Mapping
- Stochastic Feature Mapping for PAC-Bayes Classification
- Point3R: Streaming 3D Reconstruction with Explicit Spatial Pointer Memory
- Wildlife Target Re-Identification Using Self-supervised Learning in Non-Urban Settings
- On Efficient and Robust Metrics for RANSAC Hypotheses and 3D Rigid Registration
- Image-Based Alignment of 3D Scans
- Learning for mismatch removal via graph attention networks
- DeepSFM: Structure From Motion Via Deep Bundle Adjustment
- High-Precision Localization Using Ground Texture
- Active Control Points-based 6DoF Pose Tracking for Industrial Metal Objects
- Features for Ground Texture Based Localization -- A Survey
- Combinatorial clustering and the beta negative binomial process
- Deep Micro-Dictionary Learning and Coding Network
- Robust Component Detection for Flexible Manufacturing: A Deep Learning Approach to Tray-Free Object Recognition under Variable Lighting
- LoD-Loc v2: Aerial Visual Localization over Low Level-of-Detail City Models using Explicit Silhouette Alignment
- Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space
- Learning the Matching Function
- Gated Siamese Convolutional Neural Network Architecture for Human Re-Identification
- PointSSIM: A novel low dimensional resolution invariant image-to-image comparison metric
- A survey of the recent architectures of deep convolutional neural networks
- Morphological Analusis Of The Left Ventricular Eendocardial Surface Using A Bag-Of-Features Descriptor
- Towards a more robust representation of lithic industries in archaeology: a critical review of traditional approaches and modern techniques
- Li-GS: a fast 3D Gaussian reconstruction method assisted by LiDAR point clouds
- Self-Supervised Multiview Xray Matching
- Globally-Optimal Inlier Set Maximisation for Simultaneous Camera Pose and Feature Correspondence
- Morlet wavelet transform using attenuated sliding Fourier transform and kernel integral for graphic processing unit
- Recent Advances in Zero-shot Recognition
- Human Body Parts Tracking: Applications to Activity Recognition
- Introduction to the Bag of Features Paradigm for Image Classification and Retrieval
- Understanding the Fisher Vector: a multimodal part model
- Content-Based Filtering for Video Sharing Social Networks
- Single-Frame Point-Pixel Registration via Supervised Cross-Modal Feature Matching
- A Joint Model of Language and Perception for Grounded Attribute Learning
- ZeroReg3D: A Zero-shot Registration Pipeline for 3D Consecutive Histopathology Image Reconstruction
- Infrared face recognition: a literature review
- M2SFormer: Multi-Spectral and Multi-Scale Attention with Edge-Aware Difficulty Guidance for Image Forgery Localization
- TUS-REC2024: A Challenge to Reconstruct 3D Freehand Ultrasound Without External Tracker
- IF-Net: An Illumination-invariant Feature Network
- Benchmarking KAZE and MCM for Multiclass Classification
- Scene Coordinate Regression with Angle-Based Reprojection Loss for Camera Relocalization
- Backpropagation Training for Fisher Vectors within Neural Networks
- Fast entropy-regularized SDP relaxations for permutation synchronization
- Simultaneous Localization and Mapping Related Datasets: A Comprehensive Survey
- Automated Calibration of Mobile Cameras for 3D Reconstruction of Mechanical Pipes
- A critical analysis of self-supervision, or what we can learn from a single image
- The MOTIF Hand: A Robotic Hand for Multimodal Observations with Thermal, Inertial, and Force Sensors
- Data Dwarfs: A Lens Towards Fully Understanding Big Data and AI Workloads
- Recent Advances and New Guidelines on Hyperspectral and Multispectral Image Fusion
- Dense v.s. Sparse: A Comparative Study of Sampling Analysis in Scene Classification of High-Resolution Remote Sensing Imagery
- Hierarchical Modeling of Multidimensional Data in Regularly Decomposed Spaces: Synthesis and Perspective
- Robust Face Recognition with Structural Binary Gradient Patterns
- Towards the Influence of Text Quantity on Writer Retrieval
- Extremely Dense Point Correspondences using a Learned Feature Descriptor
- Sketch2code: Generating a website from a paper mockup
- Co-VisiON: Co-Visibility ReasONing on Sparse Image Sets of Indoor Scenes
- Bi-level Feature Alignment for Versatile Image Translation and Manipulation
- Flower Categorization using Deep Convolutional Neural Networks
- A Survey on Deep Learning for Localization and Mapping: Towards the Age of Spatial Machine Intelligence
- LunarLoc: Segment-Based Global Localization on the Moon
- Class Agnostic Instance-level Descriptor for Visual Instance Search
- Instance Search via Instance Level Segmentation and Feature Representation
- A Large-scale Dataset and Benchmark for Similar Trademark Retrieval
- Dense 3D Displacement Estimation for Landslide Monitoring via Fusion of TLS Point Clouds and Embedded RGB Images
- DualTHOR: A Dual-Arm Humanoid Simulation Platform for Contingency-Aware Planning
- Outsource Photo Sharing and Searching for Mobile Devices With Privacy Protection
- Unsupervised Learning from Continuous Video in a Scalable Predictive Recurrent Network
- 3D Reconstruction of Whole Stomach from Endoscope Video Using Structure-from-Motion
- Automated LoD-2 Model Reconstruction from Very-HighResolution Satellite-derived Digital Surface Model and Orthophoto
- 6-DOF Feature based LIDAR SLAM using ORB Features from Rasterized Images of 3D LIDAR Point Cloud
- A Fast Projected Fixed-Point Algorithm for Large Graph Matching
- Vision-based Human Gender Recognition: A Survey
- Randomness in Deconvolutional Networks for Visual Representation
- Detecting immune cells with label-free two-photon autofluorescence and deep learning
- Learning to Navigate for Fine-grained Classification
- Local Area Transform for Cross-Modality Correspondence Matching and Deep Scene Recognition
- Optimizing affinity-based binary hashing using auxiliary coordinates
- UltraZoom: Generating Gigapixel Images from Regular Photos
- Coarse2Fine: Two-Layer Fusion For Image Retrieval
- UAV Object Detection and Positioning in a Mining Industrial Metaverse with Custom Geo-Referenced Data
- EmbodiedPlace: Learning Mixture-of-Features with Embodied Constraints for Visual Place Recognition
- All One Needs to Know about Metaverse: A Complete Survey on Technological Singularity, Virtual Ecosystem, and Research Agenda
- Faster than Fast: Accelerating Oriented FAST Feature Detection on Low-end Embedded GPUs
- SuperPlace: The Renaissance of Classical Feature Aggregation for Visual Place Recognition in the Era of Foundation Models
- Test3R: Learning to Reconstruct 3D at Test Time
- A Survey on Deep Domain Adaptation and Tiny Object Detection Challenges, Techniques and Datasets
- FCN+RL: A Fully Convolutional Network followed by Refinement Layers to Offline Handwritten Signature Segmentation
- Privacy Leakage of SIFT Features via Deep Generative Model based Image Reconstruction
- Deep-LK for Efficient Adaptive Object Tracking
- Aggregation of binary feature descriptors for compact scene model representation in large scale structure-from-motion applications
- Image-based localization using LSTMs for structured feature correlation
- Feature Complementation Architecture for Visual Place Recognition
- Efficient Object Embedding for Spliced Image Retrieval
- A Siamese Long Short-Term Memory Architecture for Human Re-Identification
- UG2+ Track 2: A Collective Benchmark Effort for Evaluating and Advancing Image Understanding in Poor Visibility Environments
- Multimodal Machine Learning: A Survey and Taxonomy
- Bounded-Distortion Metric Learning
- Content-Based Spam Filtering on Video Sharing Social Networks
- HQFNN: A Compact Quantum-Fuzzy Neural Network for Accurate Image Classification
- Splat and Replace: 3D Reconstruction with Repetitive Elements
- Texture Classification of MR Images of the Brain in ALS using CoHOG
- O-MaMa: Learning Object Mask Matching between Egocentric and Exocentric Views
- CryoFastAR: Fast Cryo-EM Ab Initio Reconstruction Made Easy
- Homography augumented momentum constrastive learning for SAR image retrieval
- A Two-point Method for PTZ Camera Calibration in Sports
- Appearance Descriptors for Person Re-identification: a Comprehensive Review
- Toward Characteristic-Preserving Image-based Virtual Try-On Network
- Do It Yourself: Learning Semantic Correspondence from Pseudo-Labels
- Locality Preserving Markovian Transition for Instance Retrieval
- Deep Learning Reforms Image Matching: A Survey and Outlook
- Metric Learning Driven Multi-Task Structured Output Optimization for Robust Keypoint Tracking
- A general albedo recovery approach for aerial photogrammetric images through inverse rendering
- A Near-Term Quantum Computing Approach for Hard Computational Problems in Space Exploration
- AWNet: Attentive Wavelet Network for Image ISP
- Performance comparison of 3D correspondence grouping algorithm for 3D plant point clouds
- Efficient Continuous Top-k Geo-Image Search on Road Network
- Multiple Measurements and Joint Dimensionality Reduction for Large Scale Image Search with Short Vectors - Extended Version
- Visor: Privacy-Preserving Video Analytics as a Cloud Service
- Hybrid BYOL-ViT: Efficient approach to deal with small datasets
- Blurring the Line Between Structure and Learning to Optimize and Adapt\n Receptive Fields
- Learning to Fuse Local Geometric Features for 3D Rigid Data Matching
- Jointly Modeling Motion and Appearance Cues for Robust RGB-T Tracking
- Exploring Auxiliary Context: Discrete Semantic Transfer Hashing for Scalable Image Retrieval
- A Review of Computer Vision Methods in Network Security
- Adversarial Image Alignment and Interpolation
- What you need to know about the state-of-the-art computational models of object-vision: A tour through the models
- Polar Transformer Networks
- Recent Advances in Deep Learning for Object Detection
- Deep Learning-Assisted Localisation of Nanoparticles in synthetically generated two-photon microscopy images
- Learning scale-variant and scale-invariant features for deep image classification
- Do We Still Need to Work on Odometry for Autonomous Driving?
- Adapted and Oversegmenting Graphs: Application to Geometric Deep Learning
- Classifying Suspicious Content in Tor Darknet
- A survey on Kornia: an Open Source Differentiable Computer Vision Library for PyTorch
- A Comprehensive Survey on Pose-Invariant Face Recognition
- Motion Equivariance OF Event-based Camera Data with the Temporal Normalization Transform
- Novel Co-variant Feature Point Matching Based on Gaussian Mixture Model
- DSAC - Differentiable RANSAC for Camera Localization
- GPGPU Acceleration of the KAZE Image Feature Extraction Algorithm
- Improving Raw Image Storage Efficiency by Exploiting Similarity
- Scale-Aware Trident Networks for Object Detection
- Automatic Face Understanding: Recognizing Families in Photos
- Object Detection Using Deep CNNs Trained on Synthetic Images
- Attending Category Disentangled Global Context for Image Classification
- Semantic Diversity versus Visual Diversity in Visual Dictionaries
- Accelerating SfM-based Pose Estimation with Dominating Set
- Convolutional neural network architecture for geometric matching
- Privacy Prediction of Images Shared on Social Media Sites Using Deep Features
- Diverse Large-Scale ITS Dataset Created from Continuous Learning for Real-Time Vehicle Detection
- A Hajj And Umrah Location Classification System For Video Crowded Scenes
- 3D Holographic Flow Cytometry Measurements of Microalgae: Strategies for Angle Recovery in Complex Rotation Patterns
- Facing & mitigating common challenges when working with real-world data: The Data Learning Paradigm
- Robust Registration of Gaussian Mixtures for Colour Transfer
- Deeply Activated Salient Region for Instance Search
- High Five: Improving Gesture Recognition by Embracing Uncertainty
- Dense Match Summarization for Faster Two-view Estimation
- Structure from motion photogrammetry in physical geography
- Towards Recognizing New Semantic Concepts in New Visual Domains
- Link and code: Fast indexing with graphs and compact regression codes
- Overcoming Labelled Data Scarcity for Defect Classification in Scanning Tunneling Microscopy
- G4Seg: Generation for Inexact Segmentation Refinement with Diffusion Models
- Object-Level Context Modeling For Scene Classification with Context-CNN
- The Image Torque Operator for Contour Processing
- Iterative Manifold Embedding Layer Learned by Incomplete Data for Large-scale Image Retrieval
- Implicit Deformable Medical Image Registration with Learnable Kernels
- SteerPose: Simultaneous Extrinsic Camera Calibration and Matching from Articulation
- K-Median Clustering, Model-Based Compressive Sensing, and Sparse Recovery for Earth Mover Distance
- MIXER: A Principled Framework for Multimodal, Multiway Data Association
- Geometric Neural Phrase Pooling: Modeling the Spatial Co-occurrence of Neurons
- Automatic Description Generation from Images: A Survey of Models, Datasets, and Evaluation Measures
- BIRL: Benchmark on Image Registration methods with Landmark validation
- Nonlinear Intensity Underwater Sonar Image Matching Method Based on Phase Information and Deep Convolution Features
- MessyTable: Instance Association in Multiple Camera Views
- Flying Co-Stereo: Enabling Long-Range Aerial Dense Mapping via Collaborative Stereo Vision of Dynamic-Baseline
- Object Detection Networks on Convolutional Feature Maps
- Multi-Level Feature Descriptor for Robust Texture Classification via Locality-Constrained Collaborative Strategy
- Multi-scale PIIFD for Registration of Multi-source Remote Sensing Images
- Mapping erosion and deposition in an agricultural landscape: Optimization of UAV image acquisition schemes for SfM-MVS
- Robust Perspective Correction for Real-World Crack Evolution Tracking in Image-Based Structural Health Monitoring
- LiftFeat: 3D Geometry-Aware Local Feature Matching
- 6D Pose Estimation on Point Cloud Data through Prior Knowledge Integration: A Case Study in Autonomous Disassembly
- 50 Years of Automated Face Recognition
- GARLIC: GAussian Representation LearnIng for spaCe partitioning
- Cross-Modal Characterization of Thin Film MoS2 Using Generative Models
- Rooms from Motion: Un-posed Indoor 3D Object Detection as Localization and Mapping
- Descriptor Matching with Convolutional Neural Networks: a Comparison to SIFT
- MCFNet: A Multimodal Collaborative Fusion Network for Fine-Grained Semantic Classification
- Holistic Large-Scale Scene Reconstruction via Mixed Gaussian Splatting
- Vector Nonlocal Euclidean Median: Principal Bundle Captures The Nature of Patch Space
- Multi-scale Orderless Pooling of Deep Convolutional Activation Features
- Unsupervised training of keypoint-agnostic descriptors for flexible retinal image registration
- TinaFace: Strong but Simple Baseline for Face Detection
- From Images to 3D Shape Attributes
- UAVPairs: A Challenging Benchmark for Match Pair Retrieval of Large-scale UAV Images
- Efficient feature matching for UAV images based on compact GPU data scheduling
- Positive Semidefinite Metric Learning with Boosting
- BigDataBench: A Scalable and Unified Big Data and AI Benchmark Suite
- Deep Learning Multi-View Representation for Face Recognition
- Large Scale Indexing of Generic Medical Image Data using Unbiased Shallow Keypoints and Deep CNN Features
- Robust Alignment of Multi-Exposed Images with Saturated Regions
- Automated detection of sewer pipe defects in closed-circuit television images using deep learning techniques
- Deep Multi-Resolution Dictionary Learning for Histopathology Image Analysis
- Hierarchical Spatial Transformer Network
- Contact Area Detector using Cross View Projection Consistency for COVID-19 Projects
- SANSA: Unleashing the Hidden Semantics in SAM2 for Few-Shot Segmentation
- Detect-to-Retrieve: Efficient Regional Aggregation for Image Search
- HS-SLAM: A Fast and Hybrid Strategy-Based SLAM Approach for Low-Speed Autonomous Driving
- Visual Loop Closure Detection Through Deep Graph Consensus
- Visual Product Graph: Bridging Visual Products And Composite Images For End-to-End Style Recommendations
- Voronoi-based compact image descriptors: Efficient Region-of-Interest retrieval with VLAD and deep-learning-based descriptors
- Real-time Prediction of Soft Tissue Deformations Using Data-driven Nonlinear Presurgical Simulations
- CAD-SLAM: Consistency-Aware Dynamic SLAM with Dynamic-Static Decoupled Mapping
- Toward Automated Discovery of Artistic Influence
- Single-Stage 6D Object Pose Estimation
- Compact Binary Fingerprint for Image Copy Re-Ranking
- Fracking Deep Convolutional Image Descriptors
- Gradient Boundary Histograms for Action Recognition
- Investigation of Multimodal Features, Classifiers and Fusion Methods for Emotion Recognition
- The Geometry of Distributed Representations for Better Alignment, Attenuated Bias, and Improved Interpretability
- Local Descriptor for Robust Place Recognition using LiDAR Intensity
- Benchmarking 6DOF Outdoor Visual Localization in Changing Conditions
- Alchemist: Turning Public Text-to-Image Data into Generative Gold
- Automated Top View Registration of Broadcast Football Videos
- PointSIFT: A SIFT-like Network Module for 3D Point Cloud Semantic Segmentation
- Semi-Lexical Languages -- A Formal Basis for Unifying Machine Learning and Symbolic Reasoning in Computer Vision
- Leveraging the Power of Gabor Phase for Face Identification: A Block Matching Approach
- RIFT: Multi-modal Image Matching Based on Radiation-invariant Feature Transform
- Cell Detection in Microscopy Images with Deep Convolutional Neural Network and Compressed Sensing
- Spatio-spectral deep learning methods for in-vivo hyperspectral laryngeal cancer detection
- On Denoising Walking Videos for Gait Recognition
- Discrete approximations of the affine Gaussian derivative model for visual receptive fields
- Distributable Consistent Multi-Object Matching
- Autocamera Calibration for traffic surveillance cameras with wide angle lenses
- Leveraging KANs for Expedient Training of Multichannel MLPs via Preconditioning and Geometric Refinement
- Semantic Correspondence: Unified Benchmarking and a Strong Baseline
- To Glue or Not to Glue? Classical vs Learned Image Matching for Mobile Mapping Cameras to Textured Semantic 3D Building Models
- Pose Estimation of Specular and Symmetrical Objects
- A simple and effective postprocessing method for image classification
- Half-CNN: A General Framework for Whole-Image Regression
- MinkUNeXt-SI: Improving point cloud-based place recognition including spherical coordinates and LiDAR intensity
- Evaluation of Three Vision Based Object Perception Methods for a Mobile Robot
- Self-paced Learning for Weakly Supervised Evidence Discovery in Multimedia Event Search
- PawPrint: Whose Footprints Are These? Identifying Animal Individuals by Their Footprints
- Attention-Aware Age-Agnostic Visual Place Recognition
- Simultaneous Detection and Segmentation
- Affine-Gradient Based Local Binary Pattern Descriptor for Texture Classiffication
- Decoupled Geometric Parameterization and its Application in Deep Homography Estimation
- Recurrent Transformer Networks for Semantic Correspondence
- SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence
- GMatch: A Lightweight, Geometry-Constrained Keypoint Matcher for Zero-Shot 6DoF Pose Estimation in Robotic Grasp Tasks
- Visual Representations: Defining Properties and Deep Approximations
- Bayes Merging of Multiple Vocabularies for Scalable Image Retrieval
- 3D Reconstruction from Sketches
- Egocentric Action-aware Inertial Localization in Point Clouds with Vision-Language Guidance
- Automatic Handgun Detection in X-ray Images using Bag of Words Model with Selective Search
- Capturing the symptoms of malicious code in electronic documents by file's entropy signal combined with Machine learning
- Semantic Matching by Weakly Supervised 2D Point Set Registration
- Relative Camera Pose Estimation Using Convolutional Neural Networks
- Generative Hierarchical Features from Synthesizing Images
- Accurate Label Refinement From Multiannotator of Remote Sensing Data
- Understanding and Improving Kernel Local Descriptors
- On the Computation of Kantorovich-Wasserstein Distances between 2D-Histograms by Uncapacitated Minimum Cost Flows
- Beyond Text: Unveiling Privacy Vulnerabilities in Multi-modal Retrieval-Augmented Generation
- Good Practice in CNN Feature Transfer
- Fully Convolutional Neural Networks for Crowd Segmentation
- Iterative Global Mapping-Local Searching for Heterogeneous Change Detection with Unregistered Images
- Deep Architectures and Ensembles for Semantic Video Classification
- Image Retrieval based on Bag-of-Words model
- Segmentation-driven 6D Object Pose Estimation
- Recent Advances in Object Detection in the Age of Deep Convolutional Neural Networks
- Learning Cross-Spectral Point Features with Task-Oriented Training
- Simultaneous 3D Object Segmentation and 6-DOF Pose Estimation
- VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold
- PeCA: Palette Context Assisted Inference for Test-Time Paint-Bucket Colourisation on Animation Videos
- Self-supervising Fine-grained Region Similarities for Large-scale Image Localization
- Extended Affinity Propagation: Global Discovery and Local Insights
- GIE-Bench: Towards Grounded Evaluation for Text-Guided Image Editing
- Nonmyopic View Planning for Active Object Detection
- Two-view 3D Reconstruction for Food Volume Estimation
- Convolutional Neural Networks for joint object detection and pose estimation: A comparative study
- Multi-Image Semantic Matching by Mining Consistent Features
- Depth Anything with Any Prior
- Deep Weakly Supervised Positioning
- VolE: A Point-cloud Framework for Food 3D Reconstruction and Volume Estimation
- Learned Lightweight Smartphone ISP with Unpaired Data
- Discriminative Unsupervised Feature Learning with Exemplar Convolutional Neural Networks
- Large scale deduplication based on fingerprints
- RPL-UIE: Reliable Prior Learning for Underwater Image Enhancement
- Multiple neuronal populations control the eating behavior in Hydra and are responsive to microbial signals
- Egg patterns as identity signals in colonial seabirds: a comparison of four alcid species
- Bio‐Metamaterials for Mechano‐Regulation of Mesenchymal Stem Cells
- Semantics-Driven Unsupervised Learning for Monocular Depth and Ego-Motion Estimation
- Multiresolution Match Kernels for Gesture Video Classification
- Neural Inertial Odometry from Lie Events
- Using Cross-Domain Detection Loss to Infer Multi-Scale Information for Improved Tiny Head Tracking
- Hybrid Scene Compression for Visual Localization
- A Robust Real-Time Computing-based Environment Sensing System for Intelligent Vehicle
- RFVTM: A Recovery and Filtering Vertex Trichotomy Matching for Remote Sensing Image Registration
- Solving k-means on High-dimensional Big Data
- Localization of Autonomous Vehicles: Proof of Concept for A Computer Vision Approach
- Self-supervised learning for autonomous vehicles perception: A conciliation between analytical and learning methods
- Primary human neutrophils and monocytes migrate along endothelial cell boundaries to optimize search efficiency under static in vitro conditions
- RDD: Robust Feature Detector and Descriptor using Deformable Transformer
- Application-Driven Near-Data Processing for Similarity Search
- Spherical Kernel for Efficient Graph Convolution on 3D Point Clouds
- Convolution Neural Network Architecture Learning for Remote Sensing Scene Classification
- DeepID-Net: multi-stage and deformable deep convolutional neural networks for object detection
- Indoor Activity Detection and Recognition for Sport Games Analysis
- NeuGen: Amplifying the 'Neural' in Neural Radiance Fields for Domain Generalization
- Survey of Filtered Approximate Nearest Neighbor Search over the Vector-Scalar Hybrid Data
- Vision Transformers and Convolutional Neural Networks for Land Use Scene Classification
- SADet: Learning An Efficient and Accurate Pedestrian Detector
- Parallel Structure from Motion from Local Increment to Global Averaging
- Rethinking the Image Feature Biases Exhibited by Deep CNN Models
- Deep Convolutional Ranking for Multilabel Image Annotation
- A Curriculum Domain Adaptation Approach to the Semantic Segmentation of Urban Scenes
- Deep Learning-Based Robust Optical Guidance for Hypersonic Platforms
- Yum-Me
- Auto-regressive transformation for image alignment
- StabStitch++: Unsupervised Online Video Stitching with Spatiotemporal Bidirectional Warps
- DiffusionSfM: Predicting Structure and Motion via Ray Origin and Endpoint Diffusion
- Pose estimation using local structure-specific shape and appearance context
- In Search of Inliers: 3D Correspondence by Local and Global Voting
- Stomach 3D Reconstruction Based on Virtual Chromoendoscopic Image Generation
- A Comprehensive Performance Evaluation for 3D Transformation Estimation Techniques
- Multimodal Machine Learning: A Survey and Taxonomy
- 3D Face Recognition: A Survey
- Local-to-Global Self-Attention in Vision Transformers
- Cloud-based Privacy Preserving Image Storage, Sharing and Search
- A Birotation Solution for Relative Pose Problems
- OBD-Finder: Explainable Coarse-to-Fine Text-Centric Oracle Bone Duplicates Discovery
- Neural Codes for Image Retrieval
- Geometric and photometric invariant distinctive regions detection
- 3D Surface Reconstruction by Pointillism
- ORB-SLAM3: An Accurate Open-Source Library for Visual, Visual-Inertial and Multi-Map SLAM
- Argus: Metric Panoramic 3D Reconstruction for Indoor Scenes
- Leveraging Localization for Multi-camera Association
- Fast Object Detection with Latticed Multi-Scale Feature Fusion
- T-Graph: Enhancing Sparse-view Camera Pose Estimation by Pairwise Translation Graph
- One-Shot Fine-Grained Instance Retrieval
- Robust Camera Location Estimation by Convex Programming
- Generating Orthophotos from Curved Heritage Interiors Using Open-Source Photogrammetry: A Case Study of Saint Theodore Church, Cappadocia
- Stereo rectification of pushbroom satellite images by robustly estimating the fundamental matrix
- Energy-Based Geometric Multi-model Fitting
- Evaluation of Pooling Operations in Convolutional Architectures for Object Recognition
- DeepV2D: Video to Depth with Differentiable Structure from Motion
- Minimal Aspect Distortion (MAD) Mosaicing of Long Scenes
- Adversarial Soft-detection-based Aggregation Network for Image Retrieval
- Rapid Near-Neighbor Interaction of High-dimensional Data via Hierarchical Clustering
- SDGMNet: Statistic-based Dynamic Gradient Modulation for Local Descriptor Learning
- On Localizing a Camera from a Single Image
- InCaRPose: In-Cabin Relative Camera Pose Estimation Model and Dataset
- Facial Landmarks Detection by Self-Iterative Regression based Landmarks-Attention Network
- A framework for robust object multi-detection with a vote aggregation and a cascade filtering
- General Purpose (GenP) Bioimage Ensemble of Handcrafted and Learned Features with Data Augmentation
- A Two-Stage Combined Classifier in Scale Space Texture Classification
- Keypoint Density-based Region Proposal for Fine-Grained Object Detection\n and Classification using Regions with Convolutional Neural Network Features
- Learning Structural Graph Layouts and 3D Shapes for Long Span Bridges 3D Reconstruction
- Differentiable Visual Computing
- Objective comparison of particle tracking methods
- Image Matching Using Generalized Scale-Space Interest Points
- Scale Selection Properties of Generalized Scale-Space Interest Point Detectors
- A computational theory of visual receptive fields
- Interlayer and Intralayer Scale Aggregation for Scale-invariant Crowd Counting
- Beyond Sharing Weights for Deep Domain Adaptation
- FPGA-based Binocular Image Feature Extraction and Matching System
- Off-road Robotics—An Overview
- Generating diverse and representative image search results for landmarks
- Local Naive Bayes Nearest Neighbor for Image Classification
- A Unified Algorithmic Framework for Multi-Dimensional Scaling
- Privacy-Preserving Content-Based Image Retrieval in the Cloud
- General Dynamic Scene Reconstruction from Multiple View Video
- Unsupervised Multi-view Clustering by Squeezing Hybrid Knowledge from Cross View and Each View
- Geometric Models with Co-occurrence Groups
- Face Identification from Manipulated Facial Images using SIFT
- PRINS: Resistive CAM Processing in Storage
- Frame-wise Motion and Appearance for Real-time Multiple Object Tracking
- Automatic Calculation of Resolution in Lateral Cephalogram Based on Scale Mark Detection
- Classic versus deep learning approaches to address computer vision challenges
- Context-Aware Embeddings for Automatic Art Analysis
- Fast Fine-Grained Image Classification via Weakly Supervised Discriminative Localization
- SPP-Net: Deep Absolute Pose Regression with Synthetic Views
- DenseBox: Unifying Landmark Localization with End to End Object Detection
- Reliable Shot Identification for Complex Event Detection via Visual-Semantic Embedding
- A Study for Universal Adversarial Attacks on Texture Recognition
- A linear method for camera pair self-calibration and multi-view reconstruction with geometrically verified correspondences
- Indicative Image Retrieval: Turning Blackbox Learning into Grey
- Which Parts Determine the Impression of the Font?
- Patterns, predictions, and actions: A story about machine learning
- Detecting Parking Spaces in a Parcel using Satellite Images
- Multi-Modal Music Information Retrieval: Augmenting Audio-Analysis with Visual Computing for Improved Music Video Analysis
- Smooth Deformation Field-based Mismatch Removal in Real-time
- MantaMatcher: automated photographic identification of manta rays using keypoint features
- SD-6DoF-ICLK: Sparse and Deep Inverse Compositional Lucas-Kanade Algorithm on SE(3)
- Monitoring the temporal evolution of a Sicilian badland area by unmanned aerial vehicles
- Topological descriptors for 3D surface analysis
- Improved Descriptors for Patch Matching and Reconstruction
- Blur-Countering Keypoint Detection via Eigenvalue Asymmetry
- Improving Image Clustering using Sparse Text and the Wisdom of the Crowds
- Synchronisation of partial multi-matchings via non-negative factorisations
- Detection of Retinal Vascular Bifurcations by Trainable V4-Like Filters
- A Survey on 3D Reconstruction Techniques in Plant Phenotyping: From Classical Methods to Neural Radiance Fields (NeRF), 3D Gaussian Splatting (3DGS), and Beyond
- Efficient Dense Matching for Enhanced Gaussian Splatting Using AV1 Motion Vectors
- Invariant object recognition is a personalized selection of invariant features in humans, not simply explained by hierarchical feed-forward vision models
- Large-scale visual SLAM for in-the-wild videos
- Joint Optimization of Neural Radiance Fields and Continuous Camera Motion from a Monocular Video
- Activity recognition with smartphone sensors
- Delta Complexes in Digital Images. Approximating Image Object Shapes
- MP-SfM: Monocular Surface Priors for Robust Structure-from-Motion
- Fast and Robust Speckle Pattern Authentication by Scale Invariant Feature Transform algorithm in Physical Unclonable Functions
- 3DPyranet Features Fusion for Spatio-temporal Feature Learning
- Fiji: an open-source platform for biological-image analysis
- Visualization of image data from cells to organisms
- A Survey of Recent View-based 3D Model Retrieval Methods
- A Comparison of Dense Region Detectors for Image Search and Fine-Grained Classification
- Diffusion-geometric maximally stable component detection in deformable shapes
- Data Fusion of Objects Using Techniques Such as Laser Scanning, Structured Light and Photogrammetry for Cultural Heritage Applications
- SGFormer: Structure-Guided Transformer for Robust Local Feature Matching
- Object Level Deep Feature Pooling for Compact Image Representation
- Wide-Depth-Range 6D Object Pose Estimation in Space
- Multi-Temporal Aerial Image Registration Using Semantic Features
- On Recognizing Transparent Objects in Domestic Environments Using Fusion\n of Multiple Sensor Modalities
- Effective Label Propagation for Discriminative Semi-Supervised Domain Adaptation
- Applications of Machine Learning in Document Digitisation
- PCANet: A Simple Deep Learning Baseline for Image Classification?
- Who ordered this?: Exploiting implicit user tag order preferences for personalized image tagging
- Generic Instance Search and Re-identification from One Example via Attributes and Categories
- Adaptive Training of Random Mapping for Data Quantization
- Data clustering: 50 years beyond K-means
- 3D tree skeletonization from multiple images based on PyrLK optical flow
- DeepBrain: Functional Representation of Neural In-Situ Hybridization Images for Gene Ontology Classification Using Deep Convolutional Autoencoders
- A generalised feature for low level vision
- Low Cost Eye Tracking: The Current Panorama
- EdgePoint2: Compact Descriptors for Superior Efficiency and Accuracy
- Shape modeling and matching in identifying 3D protein structures
- Generalized Axiomatic Scale-Space Theory
- Scale‐Space
- A connectomic study of a petascale fragment of human cerebral cortex
- A Genealogy of Foundation Models in Remote Sensing
- Maximized Posteriori Attributes Selection from Facial Salient Landmarks for Face Recognition
- RSI-CB: A Large Scale Remote Sensing Image Classification Benchmark via Crowdsource Data
- Handwritten isolated Bangla compound character recognition: A new benchmark using a novel deep learning approach
- Efficient Similarity-aware Compression to Reduce Bit-writes in\n Non-Volatile Main Memory for Image-based Applications
- Near-duplicate video detection featuring coupled temporal and perceptual visual structures and logical inference based matching
- Interferences in match kernels
- Video Imprint
- Characterizing snow roughness: A novel low-cost portable tool based on ChArUco board and digital photography
- Segmenting Epipolar Line
- Dense Semantic 3D Map Based Long-Term Visual Localization with Hybrid Features
- Registering large volume serial-section electron microscopy image sets for neural circuit reconstruction using FFT signal whitening
- Leveraging unsupervised image registration for discovery of landmark shape descriptor
- MDSSD: Multi-scale Deconvolutional Single Shot Detector for Small Objects
- SegICP: Integrated deep semantic segmentation and pose estimation
- TRLF: An Effective Semi-fragile Watermarking Method for Tamper Detection\n and Recovery based on LWT and FNN
- Efficient Interactive Search for Geo-tagged Multimedia Data
- Memory Based Online Learning of Deep Representations from Video Streams
- An Efficient Approach for Geo-Multimedia Cross-Modal Retrieval
- Active Perception for Ambiguous Objects Classification
- LoRetta: A Foundation Model and Extensive Dataset for Global-Scale Remote Sensing Dense Image Matching
- Multiview Differential Geometry of Curves
- Label-Free Target-Domain Adaptation for Unconstrained Event-Image Feature Matching via Dual-Stage Distillation
- Complete Scene Reconstruction by Merging Images and Laser Scans
- Interactive visual analysis of contrast-enhanced ultrasound data based on small neighborhood statistics
- Feature concatenation multi-view subspace clustering
- Constrained bilinear factorization multi-view subspace clustering
- PRaDA: Projective Radial Distortion Averaging
- Learning Underwater Active Perception in Simulation
- Texture2LoD3: Enabling LoD3 Building Reconstruction With Panoramic Images
- An Accelerated Camera 3DMA Framework for Efficient Urban GNSS Multipath Estimation
- TriVoC: Efficient Voting-based Consensus Maximization for Robust Point Cloud Registration with Extreme Outlier Ratios
- Nasal Patches and Curves for Expression-Robust 3D Face Recognition
- Harnessing Multi-View Perspective of Light Fields for Low-Light Imaging
- Density‐based region search with arbitrary shape for object localisation
- Learning to detect video events from zero or very few video examples
- Robust uncalibrated stereo rectification with constrained geometric distortions (USR-CGD)
- From Past to Present: A Survey of Malicious URL Detection Techniques, Datasets and Code Repositories
- SRT3D: A Sparse Region-Based 3D Object Tracking Approach for the Real World
- Artefact removal in ground truth deficient fluctuations-based nanoscopy images using deep learning
- Deep triplet hashing network for case-based medical image retrieval
- Weakly Supervised PatchNets: Describing and Aggregating Local Patches for Scene Recognition
- A Functional Representation for Graph Matching
- Semantic mapping for orchard environments by merging two‐sides reconstructions of tree rows
- Quasi-homography warps in image stitching
- Joint Forward-Backward Visual Odometry for Stereo Cameras
- From Volcano to Toyshop
- Fully automatic segmentation and objective assessment of atrial scars for long‐standing persistent atrial fibrillation patients using late gadolinium‐enhanced MRI
- Matching Algorithms: Fundamentals, Applications and Challenges
- Multi-view metric learning for multi-instance image classification
- Object-Based Visual Camera Pose Estimation From Ellipsoidal Model and 3D-Aware Ellipse Prediction
- Feature Identification and Matching for Hand Hygiene Pose
- Coarse-to-Fine Registration of Airborne LiDAR Data and Optical Imagery on Urban Scenes
- Nazr-CNN: Fine-Grained Classification of UAV Imagery for Damage Assessment
- A Review of Visual Trackers and Analysis of its Application to Mobile Robot
- DeepMorph: A System for Hiding Bitstrings in Morphable Vector Drawings
- Deep learning-guided surface characterization for autonomous hydrogen lithography
- SAGA: Semantic-Aware Gray color Augmentation for Visible-to-Thermal Domain Adaptation across Multi-View Drone and Ground-Based Vision Systems
- Pose Optimization for Autonomous Driving Datasets using Neural Rendering Models
- Scalable High-Precision Microfabrication on Various Lithography-Incompatible Substrates and Materials Enabled by Wafer-Scale Transfer Lithography of Commercial Photoresists
- Locally Supervised Deep Hybrid Model for Scene Recognition
- MS-POFT: multiscale phase-orientation guided feature transform for multi-modal image matching
- Performance Analysis and Robustification of Single-query 6-DoF Camera Pose Estimation
- Perception and Navigation in Autonomous Systems in the Era of Learning: A Survey
- RISAS: A Novel Rotation, Illumination, Scale Invariant Appearance and Shape Feature
- Text Detection and Recognition in the Wild: A Review
- Depth-Aware Multi-Grid Deep Homography Estimation With Contextual Correlation
- Recovering Spatiotemporal Correspondence between Deformable Objects by Exploiting Consistent Foreground Motion in Video
- An Improved Observation Model for Super-Resolution Under Affine Motion
- Registration of Sub-Sequence and Multi-Camera Reconstructions for Camera Motion Estimation
- Invariant properties of a locally salient dither pattern with a spatial-chromatic histogram
- Behavior Discovery and Alignment of Articulated Object Classes from Unstructured Video
- Exploring Generalizable Pre-training for Real-world Change Detection via Geometric Estimation
- Fast Exact Search in Hamming Space with Multi-Index Hashing
- To Learn or Not to Learn: Visual Localization from Essential Matrices
- Fast Transport Optimization for Monge Costs on the Circle
- Fast Loop Closure Detection via Binary Content
- A Spatial Layout and Scale Invariant Feature Representation for Indoor Scene Classification
- Network Uncertainty Informed Semantic Feature Selection for Visual SLAM
- Joint Facade Registration and Segmentation for Urban Localization
- Deep learning for video classification and captioning
- Res2Net: A New Multi-Scale Backbone Architecture
- Land Use Classification in Remote Sensing Images by Convolutional Neural Networks
- Discriminative Nonlinear Analysis Operator Learning: When Cosparse Model Meets Image Classification
- PGD-UNet: A Position-Guided Deformable Network for Simultaneous Segmentation of Organs and Tumors
- An empirical study on the effects of different types of noise in image classification tasks
- Deep divergence-based approach to clustering
- A PCB Dataset for Defects Detection and Classification
- CARE: Content Aware Redundancy Elimination for Disaster Communications\n on Damaged Networks
- Two-view correspondence learning using graph neural network with reciprocal neighbor attention
- Active Testing for Face Detection and Localization
- Robust Registration of Multimodal Remote Sensing Images Based on Structural Similarity
- On Image segmentation using Fractional Gradients-Learning Model Parameters using Approximate Marginal Inference
- A Fast Keypoint Based Hybrid Method for Copy Move Forgery Detection
- Detection and Recognition of Malaysian Special License Plate Based On SIFT Features
- Texture classification using block intensity and gradient difference (BIGD) descriptor
- Deep Learning for Scene Classification: A Survey
- Fake Generated Painting Detection Via Frequency Analysis
- What Looks Good with my Sofa: Ensemble Multimodal Search for Interior Design
- Scale Adaptive Clustering of Multiple Structures
- Video Summarization using Deep Semantic Features
- Efficient Non-Consecutive Feature Tracking for Robust Structure-From-Motion
- Object-sensitive Deep Reinforcement Learning
- WarpNet: Weakly Supervised Matching for Single-view Reconstruction
- Fast k Nearest Neighbor Search using GPU
- New bag of deep visual words based features to classify chest x-ray images for COVID-19 diagnosis
- High-quality Panorama Stitching based on Asymmetric Bidirectional Optical Flow
- Registration of Images With Outliers Using Joint Saliency Map
- Structure-aware completion of photogrammetric meshes in urban road environment
- AerialMegaDepth: Learning Aerial-Ground Reconstruction and View Synthesis
- Regist3R: Incremental Registration with Stereo Foundation Model
- A Multi-Layer System for Ultra-High-Resolution Static 360-Degree Telepresence
- BendTwin: Robust Dense-to-Sparse Physical Reconstruction with Bending-Aware Differentiable Spring-Mass Models
- How Smart Does Your Profile Image Look?
- MirBot: A collaborative object recognition system for smartphones using convolutional neural networks
- Retinal Imaging and Image Analysis
- Are Pretrained Image Matchers Good Enough for SAR-Optical Satellite Registration?
- Exploring structure for long-term tracking of multiple objects in sports videos
- An Image is Worth K Topics: A Visual Structural Topic Model with Pretrained Image Embeddings
- SeeTree -- A modular, open-source system for tree detection and orchard localization
- Comparing Performance of Preprocessing Techniques for Traffic Sign Recognition Using a HOG-SVM
- Knowledge Distillation for Underwater Feature Extraction and Matching via GAN-synthesized Images
- Hardware, Algorithms, and Applications of the Neuromorphic Vision Sensor: a Review
- LookingGlass: Generative Anamorphoses via Laplacian Pyramid Warping
- Parallax Effect Free Mosaicing of Underwater Video Sequence Based on Texture Features
- TokenFocus-VQA: Enhancing Text-to-Image Alignment with Position-Aware Focus and Multi-Perspective Aggregations on LVLMs
- Novel Diffusion Models for Multimodal 3D Hand Trajectory Prediction
- To Match or Not to Match: Revisiting Image Matching for Reliable Visual Place Recognition
- Learning Affine Correspondences by Integrating Geometric Constraints