2024/12/09 by Nicolas Dufour, David Picard, Dufour, Nicolas +6 · 2 voices · 12 citations
Social Sciences · #Artificial intelligence #Computer graphics (images) #Computer science #Computer vision #Generative grammar #Generative model #Geographic Information Systems Studies #Geolocation #World Wide Web
paper · pdf · doi:10.48550/arxiv.2412.06781
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/12/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Global visual geolocation predicts where an image was captured on Earth. Since images vary in how precisely they can be localized, this task inherently involves a significant degree of ambiguity. However, existing approaches are deterministic and overlook this aspect. In this paper, we aim to close the gap between traditional geolocalization and modern generative methods. We propose the first generative geolocation approach based on diffusion and Riemannian flow matching, where the denoising process operates directly on the Earth's surface. Our model achieves state-of-the-art performance on three visual geolocation benchmarks: OpenStreetView-5M, YFCC-100M, and iNat21. In addition, we introduce the task of probabilistic visual geolocation, where the model predicts a probability distribution over all possible locations instead of a single point. We introduce new metrics and baselines for this task, demonstrating the advantages of our diffusion-based approach. Codes and models will be made available.