DINOv3
2025/08/13 by Siméoni, Oriane, Vo, Huy V., Seitzer, Maximilian +23 · 241 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
paper · doi:10.48550/arxiv.2508.10104
Abstract
Self-supervised learning holds the promise of eliminating the need for manual data annotation, enabling models to scale effortlessly to massive datasets and larger architectures. By not being tailored to specific tasks or domains, this training paradigm has the potential to learn visual representations from diverse sources, ranging from natural to aerial images -- using a single algorithm. This technical report introduces DINOv3, a major milestone toward realizing this vision by leveraging simple yet effective strategies. First, we leverage the benefit of scaling both dataset and model size by careful data preparation, design, and optimization. Second, we introduce a new method called Gram anchoring, which effectively addresses the known yet unsolved issue of dense feature maps degrading during long training schedules. Finally, we apply post-hoc strategies that further enhance our models' flexibility with respect to resolution, model size, and alignment with text. As a result, we present a versatile vision foundation model that outperforms the specialized state of the art across a broad range of settings, without fine-tuning. DINOv3 produces high-quality dense features that achieve outstanding performance on various vision tasks, significantly surpassing previous self- and weakly-supervised foundation models. We also share the DINOv3 suite of vision models, designed to advance the state of the art on a wide spectrum of tasks and data by providing scalable solutions for diverse resource constraints and deployment scenarios.
Cited by
- Toward Stable Semi-Supervised Remote Sensing Segmentation via Co-Guidance and Co-Fusion
- Retrieve and Segment: Are a Few Examples Enough to Bridge the Supervision Gap in Open-Vocabulary Segmentation?
- ReDiF: Reinforced Distillation for Few Step Diffusion
- FinPercep-RM: A Fine-grained Reward Model and Co-evolutionary Curriculum for RL-based Real-world Super-Resolution
- RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
- Are Vision Foundation Models Foundational for Electron Microscopy Image Segmentation?
- A satellite foundation model for improved wealth monitoring
- Learning Traversability-Aware Global Planners for Long Horizon Off-Road Navigation
- UMI3D: Robust 3D Generation on Unconstrained Multi-Image Inputs via Simultaneous Focus Cross-Attention Routing
- ESRVS: Extreme Semi-Supervised Retinal Vessel Segmentation with a Single Annotated Image
- DailyBench: A Unified Benchmark for AI-Generated and Manipulated Images from Modern Generative Models
- Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
- Robustifying pathology foundation models via fine-tuning
- ChangeFlow -- Latent Rectified Flow for Change Detection in Remote Sensing
- Personalize Your Large Vision-language Models With In-context Prompt Tuning
- Match-Any-Events: Zero-Shot Motion-Robust Feature Matching Across Wide Baselines for Event Cameras
- Self-Distillation of Hidden Layers for Self-Supervised Representation Learning
- EraseLoRA: MLLM-Driven Foreground Exclusion and Background Subtype Aggregation for Dataset-Free Object Removal
- Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding
- Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
- Beyond Weight Adaptation: Feature-Space Domain Injection for Cross-Modal Ship Re-Identification
- CHAMMI-75: Pre-training multi-channel models with heterogeneous microscopy images
- AlignPose: Generalizable 6D Pose Estimation via Multi-view Feature-metric Alignment
- Probing and Leveraging Video Diffusion Transformer Features for Robust Point Tracking
- SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
- SemanticGen: Video Generation in Semantic Space
- LoGoPlanner: Localization Grounded Navigation Policy with Metric-aware Visual Geometry
- UniMPR: A Unified Framework for Multimodal Place Recognition with Heterogeneous Sensor Configurations
- SG-RIFE: Semantic-Guided Real-Time Intermediate Flow Estimation with Diffusion-Competitive Perceptual Quality
- RecurGS: Interactive Scene Modeling via Discrete-State Recurrent Gaussian Fusion
- Unifying Deep Predicate Invention with Pre-trained Foundation Models
- MMLANDMARKS: a Cross-View Instance-Level Benchmark for Geo-Spatial Understanding
- DVGT: Driving Visual Geometry Transformer
- Depth Any Panoramas: A Foundation Model for Panoramic Depth Estimation
- OPENTOUCH: Bringing Full-Hand Touch to Real-World Interaction
- Radiology Report Generation with Layer-Wise Anatomical Attention
- SARMAE: Masked Autoencoder for SAR Representation Learning
- Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models
- Sceniris: A Fast Procedural Scene Generation Framework
- Multi-View Foundation Models
- MoonSeg3R: Monocular Online Zero-Shot Segment Anything in 3D with Reconstructive Foundation Priors
- BEV-Patch-PF: Particle Filtering with BEV-Aerial Feature Matching for Off-Road Geo-Localization
- Native and Compact Structured Latents for 3D Generation
- GaussianPlant: Structure-aligned Gaussian Splatting for 3D Reconstruction of Plants
- World Models Can Leverage Human Videos for Dexterous Manipulation
- Revisiting 2D Foundation Models for Scalable 3D Medical Image Classification
- Unlocking Generalization in Polyp Segmentation with DINO Self-Attention "keys"
- JoVA: Unified Multimodal Learning for Joint Video-Audio Generation
- Patch-wise Retrieval: A Bag of Practical Techniques for Instance-level Matching
- Ultra-Low Bitrate Perceptual Image Compression with Shallow Encoder
- V-Warper: Appearance-Consistent Video Diffusion Personalization via Value Warping
- RePack then Refine: Efficient Diffusion Transformer with Vision Foundation Model
- Brain-Semantoks: Learning Semantic Tokens of Brain Dynamics with a Self-Distilled Foundation Model
- FreqDINO: Frequency-Guided Adaptation for Generalized Boundary-Aware Ultrasound Image Segmentation
- Fast-FoundationStereo: Real-Time Zero-Shot Stereo Matching
- E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-training
- SoccerMaster: A Vision Foundation Model for Soccer Understanding
- Geo6DPose: Fast Zero-Shot 6D Object Pose Estimation via Geometry-Filtered Feature Matching
- YOPO-Nav: Visual Navigation using 3DGS Graphs from One-Pass Videos
- IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting
- UnReflectAnything: RGB-Only Highlight Removal by Rendering Synthetic Specular Supervision
- A Multi-Robot Platform for Robotic Triage Combining Onboard Sensing and Foundation Models
- Photo3D: Advancing Photorealistic 3D Generation through Structure-Aligned Detail Enhancement
- DINO-BOLDNet: A DINOv3-Guided Multi-Slice Attention Network for T1-to-BOLD Generation
- ConceptPose: Training-Free Zero-Shot Object Pose Estimation using Concept Vectors
- Selfi: Self Improving Reconstruction Engine via 3D Geometric Feature Alignment
- Is Generation Required for Data-Efficient Perception?
- TrajMoE: Scene-Adaptive Trajectory Planning with Mixture of Experts and Reinforcement Learning
- From Orbit to Ground: Generative City Photogrammetry from Extreme Off-Nadir Satellite Images
- Multi-view Pyramid Transformer: Look Coarser to See Broader
- See in Depth: Training-Free Surgical Scene Segmentation with Monocular Depth Priors
- SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World Model
- Few-Shot Open-Vocabulary Remote Sensing Segmentation via Textual Inversion
- Coordinated Humanoid Manipulation with Choice Policies
- I2I-Bench: A Comprehensive Benchmark Suite for Image-to-Image Editing Models
- DeRA: Decoupled Representation Alignment for Video Tokenization
- C3G: Learning Compact 3D Representations with 2K Gaussians
- ShelfGaussian: Shelf-Supervised Open-Vocabulary Gaussian-based 3D Scene Understanding
- Emergent Outlier View Rejection in Visual Geometry Grounded Transformers
- CAMEO: Correspondence-Attention Alignment for Multi-View Diffusion Models
- TEXTRIX: Latent Attribute Grid for Native Texture Generation and Beyond
- Polar Perspectives: Evaluating 2-D LiDAR Projections for Robust Place Recognition with Visual Foundation Models
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- MV-TAP: Tracking Any Point in Multi-View Videos
- KM-ViPE: Online Tightly Coupled Vision-Language-Geometry Fusion for Open-Vocabulary Semantic SLAM
- Bridging the Scale Gap: Balanced Tiny and General Object Detection in Remote Sensing Imagery
- Lost in Distortion: Uncovering the Domain Gap Between Computer Vision and Brain Imaging -- A Study on Pretraining for Age Prediction
- Goal-Driven Reward by Video Diffusion Models for Reinforcement Learning
- Mammo-FM: Breast-specific foundational model for Integrated Mammographic Diagnosis, Prognosis, and Reporting
- RobotSeg: A Model and Dataset for Segmenting Robots in Image and Video
- Contrastive Heliophysical Image Pretraining for Solar Dynamics Observatory Records
- TraceGen: World Modeling in 3D Trace Space Enables Learning from Cross-Embodiment Videos
- Seeing without Pixels: Perception from Camera Trajectories
- Cleaning the Pool: Progressive Filtering of Unlabeled Pools in Deep Active Learning
- EoS-FM: Can an Ensemble of Specialist Models act as a Generalist Feature Extractor?
- Uni-Hema: Unified Model for Digital Hematopathology
- Semantic-Enhanced Feature Matching with Learnable Geometric Verification for Cross-Modal Neuron Registration
- BotaCLIP: Contrastive Learning for Botany-Aware Representation of Earth Observation Data
- DINO-Tok: Adapting DINO for Visual Tokenizers
- 3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene Understanding
- Wanderland: Geometrically Grounded Simulation for Open-World Embodied AI
- TaCo: Capturing Spatio-Temporal Semantic Consistency in Remote Sensing Change Detection
- PromptMoG: Enhancing Diversity in Long-Prompt Image Generation via Prompt Embedding Mixture-of-Gaussian Sampling
- OmniRefiner: Reinforcement-Guided Local Diffusion Refinement
- Supervise Less, See More: Training-free Nuclear Instance Segmentation with Prototype-Guided Prompting
- NAF: Zero-Shot Feature Upsampling via Neighborhood Attention Filtering
- stable-pretraining-v1: Foundation Model Research Made Simple
- Muskie: Multi-view Masked Image Modeling for 3D Vision Pre-training
- SPIDER: Spatial Image CorresponDence Estimator for Robust Calibration
- DReX: Pure Vision Fusion of Self-Supervised and Convolutional Representations for Image Complexity Prediction
- Align & Invert: Solving Inverse Problems with Diffusion and Flow-based Models via Representation Alignment
- Towards Open-Ended Visual Scientific Discovery with Sparse Autoencoders
- Weakly Supervised Segmentation and Classification of Alpha-Synuclein Aggregates in Brightfield Midbrain Images
- VIRAL: Visual Sim-to-Real at Scale for Humanoid Loco-Manipulation
- UniSER: A Foundation Model for Unified Soft Effects Removal
- RoMa v2: Harder Better Faster Denser Feature Matching
- Learning to See Through a Baby's Eyes: Early Visual Diets Enable Robust Visual Intelligence in Humans and Machines
- Find the Leak, Fix the Split: Cluster-Based Method to Prevent Leakage in Video-Derived Datasets
- Tissue Aware Nuclei Detection and Classification Model for Histopathology Images
- Pixels or Positions? Benchmarking Modalities in Group Activity Recognition
- DINO-Detect: A Simple yet Effective Framework for Blur-Robust AI-Generated Image Detection
- DINOv3-Guided Cross Fusion Framework for Semantic-aware CT generation from MRI and CBCT
- The Persistence of Cultural Memory: Investigating Multimodal Iconicity in Diffusion Models
- Sat2RealCity: Geometry-Aware and Appearance-Controllable 3D Urban Generation from Satellite Imagery
- Φeat: Physically Grounded Material Feature Representation
- STORM: Segment, Track, and Object Re-Localization from a Single Image
- RF-DETR: Neural Architecture Search for Real-Time Detection Transformers
- Measuring the Intrinsic Dimension of Earth Representations
- Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention
- Clifford Algebraic Rotor Embeddings : Maybe embeddings should start to CARE
- DWFF-Net : A Multi-Scale Farmland System Habitat Identification Method with Adaptive Dynamic Weight
- A Circular Argument : Does RoPE need to be Equivariant for Vision?
- UniADC: A Unified Framework for Anomaly Detection and Classification
- PuzLM: Solving Jigsaw Puzzles with Sequence-to-Sequence Language Models
- VLM-driven Skill Selection for Robotic Assembly Tasks
- SA-EMO: Structure-Aligned Encoder Mixture of Operators for Generalizable Full-waveform Inversion
- From Linear Probing to Joint-Weighted Token Hierarchy: A Foundation Model Bridging Global and Cellular Representations in Biomarker Detection
- UniSplat: Unified Spatio-Temporal Fusion via 3D Latent Scaffolds for Dynamic Driving Scene Reconstruction
- GeoCrossBench: Cross-Band Generalization for Remote Sensing
- Challenging DINOv3 Foundation Model under Low Inter-Class Variability: A Case Study on Fetal Brain Ultrasound
- Image Hashing via Cross-View Code Alignment in the Age of Foundation Models
- MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-Experts
- FaCT: Faithful Concept Traces for Explaining Neural Network Decisions
- Spatially Grounded Concept Bottleneck Models via Part-Factorized Attention
- Enfold: Folding World-Generator Computation into Predictive Representations for Efficient Embodied Control
- JEPADepth: Masked Predictive Representation Learning for Self-Supervised Monocular Depth Estimation
- Representation Trajectories Matters: Complementary Evidence for OOD Detection and Image Classification
- Lightweight Image Classification of Raptor Species for Edge Devices: Rare-Species Dataset Expansion via Video Frame Extraction, Knowledge Distillation, and TensorRT Deployment
- Rad-JEPA 3D: Radiology Joint-Embedding Predictive Model for 3D Computed Tomography
- Face-Trace: Open-Set Attribution and Progressive Discovery of Synthetic Face Generators
- Robust RPC Bundle Adjustment for Multi-Date Satellite Imagery with Season-Invariant Correspondences
- Is Dimensionality a Barrier for Retrieval Models?
- SPROUT: A Scalable Diffusion Foundation Model for Agricultural Vision
- Does Object Binding Naturally Emerge in Large Pretrained Vision Transformers?
- FreeFuse: Multi-Subject LoRA Fusion via Adaptive Token-Level Routing at Test Time
- Understanding What Is Not Said:Referring Remote Sensing Image Segmentation with Scarce Expressions
- IGGT: Instance-Grounded Geometry Transformer for Semantic 3D Reconstruction
- CogStereo: Neural Stereo Matching with Implicit Spatial Cognition Embedding
- FlowOpt: Fast Optimization Through Whole Flow Processes for Training-Free Editing
- S3OD: Towards Generalizable Salient Object Detection with Synthetic Data
- Data-Centric Lessons To Improve Speech-Language Pretraining
- UWBench: A Comprehensive Vision-Language Benchmark for Underwater Understanding
- VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
- Elastic ViTs from Pretrained Models without Retraining
- Exploring Structural Degradation in Dense Representations for Self-supervised Learning
- Universal and Transferable Attacks on Pathology Foundation Models
- Memory-SAM: Human-Prompt-Free Tongue Segmentation via Retrieval-to-Prompt
- Comprehensive language-image pre-training for 3D medical image understanding
- Vision-Centric Activation and Coordination for Multimodal Large Language Models
- MUSE: Model-based Uncertainty-aware Similarity Estimation for zero-shot 2D Object Detection and Segmentation
- Rethinking the Simulation vs. Rendering Dichotomy: No Free Lunch in Spatial World Modelling
- Scope: Selective Cross-modal Orchestration of Visual Perception Experts
- AnyUp: Universal Feature Upsampling
- DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving
- UniCoD: Enhancing Robot Policy via Unified Continuous and Discrete Representation Learning
- FRIEREN: Federated Learning with Vision-Language Regularization for Segmentation
- Vision4PPG: Emergent PPG Analysis Capability of Vision Foundation Models for Vital Signs like Blood Pressure
- SpaceVista: All-Scale Visual Spatial Reasoning from mm to km
- FreeViS: Training-free Video Stylization with Inconsistent References
- Vision Language Models: A Survey of 26K Papers
- XYZCylinder: Towards Compatible Feed-Forward 3D Gaussian Splatting for Driving Scenes via Unified Cylinder Lifting Method
- Unlocking 3D Affordance Segmentation with 2D Semantic Knowledge
- Vi-TacMan: Articulated Object Manipulation via Vision and Touch
- OneVision: An End-to-End Generative Framework for Multi-view E-commerce Vision Search
- VENTURA: Adapting Image Diffusion Models for Unified Task Conditioned Navigation
- TFM Dataset: A Novel Multi-task Dataset and Integrated Pipeline for Automated Tear Film Break-Up Segmentation
- Boomerang Distillation Enables Zero-Shot Model Size Interpolation
- Activation Quantization of Vision Encoders Needs Prefixing Registers
- Domain Generalization Under Posterior Drift
- Exploring Instruction Data Quality for Explainable Image Quality Assessment
- Geometry Meets Vision: Revisiting Pretrained Semantics in Distilled Fields
- SoftCFG: Uncertainty-guided Stable Guidance for Visual Autoregressive Model
- Batch-CAM: Introduction to better reasoning in convolutional deep learning models
- CardioBench: Do Echocardiography Foundation Models Generalize Beyond the Lab?
- PatchEAD: Unifying Industrial Visual Prompting Frameworks for Patch-Exclusive Anomaly Detection
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- UniLat3D: Geometry-Appearance Unified Latents for Single-Stage 3D Generation
- AnyDepth: Depth Estimation Made Easy
- PartCo: Part-Level Correspondence Priors Enhance Category Discovery
- Spectral Collapse Drives Loss of Plasticity in Deep Continual Learning
- Johnson-Lindenstrauss Lemma Guided Network for Efficient 3D Medical Segmentation
- On the Status of Foundation Models for SAR Imagery
- SiNGER: A Clearer Voice Distills Vision Transformers Further
- Real-Time Object Detection Meets DINOv3
- Foundation Models for Face Presentation Attack Detection: A Unified Linear-Probing Benchmark
- MeshFM: 2D Features Are All You Need for 3D Shape Understanding
- SiamJEPA: On the Role of Siamese Student Encoders in JEPA
- Towards Real-Time PixOOD: Efficient Anomaly Segmentation for Autonomous Vehicles
- A scalar per patch from pre-trained ViTs enables fast moving navigation in the real world
- VFM-Loc: Training-Free Cross-View Geo-Localization via Aligning Discriminative Visual Hierarchies
- Temporal Straightening for Latent Planning
- Reading Images Like Texts: Sequential Image Understanding in Vision-Language Models
- Revisiting Vision Language Foundations for No-Reference Image Quality Assessment
- DINOv3-Diffusion Policy: Self-Supervised Large Visual Model for Visuomotor Diffusion Policy Learning
- MRN: Harnessing 2D Vision Foundation Models for Diagnosing Parkinson's Disease with Limited 3D MR Data
- DREAM: Domain-aware Reasoning for Efficient Autonomous Underwater Monitoring
- EdiVal-Agent: An Object-Centric Framework for Automated, Fine-Grained Evaluation of Multi-Turn Editing
- Brought a Gun to a Knife Fight: Modern VFM Baselines Outgun Specialized Detectors on In-the-Wild AI Image Detection
- Mars Traversability Prediction: A Multi-modal Self-supervised Approach for Costmap Generation
- Beyond Instance Consistency: Investigating View Diversity in Self-supervised Learning
- PeftCD: Leveraging Vision Foundation Models with Parameter-Efficient Fine-Tuning for Remote Sensing Change Detection
- Rethinking the Backbone in Class Imbalanced Federated Source Free Domain Adaptation: The Utility of Vision Foundation Models
- RINO: Renormalization Group Invariance with No Labels
- Understanding Ice Crystal Habit Diversity with Self-Supervised Learning
- Does DINOv3 Set a New Medical Vision Standard? Benchmarking 2D and 3D Classification, Segmentation, and Registration
- UniVerse-1: Unified Audio-Video Generation via Stitching of Experts
- Patch-Level Kernel Alignment for Dense Self-Supervised Learning
- Generalist versus Specialist Vision Foundation Models for Ocular Disease and Oculomics
- FastVGGT: Training-Free Acceleration of Visual Geometry Transformer
- DIET-CP: Lightweight and Data Efficient Self Supervised Continued Pretraining
- M3Ret: Unleashing Zero-shot Multimodal Medical Image Retrieval via Self-Supervision
- SegDINO: An Efficient Design for Medical and Natural Image Segmentation with DINO-V3
- Webly-Supervised Image Manipulation Localization via Category-Aware Auto-Annotation
- Dino U-Net: Exploiting High-Fidelity Dense Features from Foundation Models for Medical Image Segmentation
- The point is the mask: scaling coral reef segmentation with weak supervision
- Annotation-Free Open-Vocabulary Segmentation for Remote-Sensing Images
- HiRQA: Hierarchical Ranking and Quality Alignment for Opinion-Unaware Image Quality Assessment
- DINOv3 with Test-Time Training for Medical Image Registration
- VFM-Guided Semi-Supervised Detection Transformer under Source-Free Constraints for Remote Sensing Object Detection
- MAESTRO: Masked AutoEncoders for Multimodal, Multitemporal, and Multispectral Earth Observation Data
- OMGSR: You Only Need One Mid-timestep Guidance for Real-World Image Super-Resolution
- Deepfake Detection that Generalizes Across Benchmarks
- From Seeing to Experiencing: Scaling Navigation Foundation Models with Reinforcement Learning
- Image Generators are Generalist Vision Learners
- Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
- FISHER: A Foundation Model for Multi-Modal Industrial Signal Comprehensive Representation
- Towards channel foundation models (CFMs): Motivations, methodologies and opportunities
- UniLGL: Learning Uniform Place Recognition for FOV-limited/Panoramic LiDAR Global Localization
- Towards Depth Foundation Model: Recent Trends in Vision-Based Depth Estimation
- Self-supervised pretraining of vision transformers for animal behavioral analysis and neural encoding
Related