DINOv2: Learning Robust Visual Features without Supervision
2023/04/14 by Maxime Oquab, Oquab, Maxime, Timothée Darcet +50 · 5 voices · 1804 citations
Computer Science · #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning #Advanced Neural Network Applications
paper · pdf · doi:10.48550/arxiv.2304.07193
Abstract
The recent breakthroughs in natural language processing for model pretraining on large quantities of data have opened the way for similar foundation models in computer vision. These models could greatly simplify the use of images in any system by producing all-purpose visual features, i.e., features that work across image distributions and tasks without finetuning. This work shows that existing pretraining methods, especially self-supervised methods, can produce such features if trained on enough curated data from diverse sources. We revisit existing approaches and combine different techniques to scale our pretraining in terms of data and model size. Most of the technical contributions aim at accelerating and stabilizing the training at scale. In terms of data, we propose an automatic pipeline to build a dedicated, diverse, and curated image dataset instead of uncurated data, as typically done in the self-supervised literature. In terms of models, we train a ViT model (Dosovitskiy et al., 2020) with 1B parameters and distill it into a series of smaller models that surpass the best available all-purpose features, OpenCLIP (Ilharco et al., 2021) on most of the benchmarks at image and pixel levels.
Citations
Cited by
- SM4RT: Learning Structured Motion Geometry for 4D Reconstruction
- Twins: Learn to Predict Unified Representations with Focal Loss
- CorVS+: Correspondence-Driven Association of Video Trajectories and Sensors for Identity-Aware Person Localization in Warehouses
- Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability
- Atlas 2 -- Foundation models for clinical deployment
- SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation
- dRAE: Representation Autoencoder with Hyper-Spherical Codes
- Unboxing Diffusion Models for the Arts: Interactive Model Bending and Practice-Based Explainability
- Persistent Computational State: A Session-Centric Runtime for Generative World Models
- xperception -- Making Robotic Grasping Easier
- Depth-Regularized JEPA World Models Learn More Transferable Representations from Real Outdoor Robot Data
- SeeSE3: Emergence of 3D Space in Vision Features
- AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching
- Extending a Large View Synthesis Model for Multi-view Panoptic Segmentation
- Conditioning Residuals for Diffusion Models via Representation Feedback
- Real-Time Human Reconstruction and Animation using Feed-Forward Gaussian Splatting
- DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation
- Back to Back with a Copy: A Computational Analysis of AI-Generated Visual Contemporary Art Pastiches
- Development of an automated, reliable, and clinically meaningful artificial intelligence (AI) tool for diagnosing cardiac disease from conventional cardiovascular magnetic resonance (CMR) images
- LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition
- ReferTrack: Referring Then Tracking for Embodied Visual Tracking
- Trusting What You Cannot See: Auditable Fine-Tuning and Inference for Proprietary AI
- Dominant vs. Dominated: Concept-Level Generative Collapse in Diffusion Models
- IConE: Batch Independent Collapse Prevention for Self-Supervised Representation Learning
- Pixel-Space Diffusion Transformers
- Learning Semantic-Robust Change Detection via Semantic-Invariant Self-Distillation
- Mitigating Modality and Language-Style Gaps for Zero-Shot Video Moment Retrieval
- LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs
- DifFoundMAD: Foundation Models meet Differential Morphing Attack Detection
- ATSplat: Compact Feed-forward 3D Gaussian Splatting with Adaptive Token Expansion
- Show Me Examples: Inferring Visual Concepts from Image Sets
- Post-Training in Time Series Foundation Models: A Unifying Framework
- HistoFID- Calibrating Frechet-distance evaluation across pathology foundation models
- Axolotl3D: a Unified Framework for Faithful 3D Shape Completion
- FELT: Generating Tactile Signals from Vision for Visuo-Tactile Manipulation
- Evaluating Uncertainty and Quality of Visual Language Action-enabled Robots
- RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation
- Strength-Parity Ensembling with Parameter-Isolated Experts for Multi-Task Affect Recognition
- Bayesian uncertainty estimation improves clinical decision making in medical AI agents
- GigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Whole-Slide and Tumor Microenvironment Analysis
- RoboInspector: Unveiling the Unreliability of Policy Code for LLM-enabled Robotic Manipulation
- VQ-Transplant: Efficient VQ-Module Integration for Pre-trained Visual Tokenizers
- The pretraining domain outweighs the training objective in setting the privacy-utility trade-off of differentially private medical image analysis
- VG3S: Visual Geometry Grounded Gaussian Splatting for Semantic Occupancy Prediction
- To Blend In, First Decouple: Rethinking Camouflage Image Generation via Context-Decoupled Representations
- Brewing Stronger Features: Dual-Teacher Distillation for Multispectral Earth Observation
- MindPilot: Closed-loop Visual Stimulation Optimization for Brain Modulation with EEG-guided Diffusion
- MATANet: A Multi-context Attention and Taxonomy-Aware Network for Fine-Grained Underwater Recognition of Marine Species
- CNS-Edit++: Category-Agnostic 3D Editing with Coupled Neural Shape Representation
- HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement
- Test-Time Registers as Global Priors for Tokenized Image Generation
- CONTACT: CONtact-aware TACTile Learning for Robotic Disassembly
- Zero-Shot DINOv3-Based Image Matching via Many-to-Many Association
- Lifting Embodied World Models for Planning and Control
- In-Context Learning for Wound Classification with Small Multimodal Language Models
- Norm or Direction? Decoding Vision Mambas for High-Resolution Vision
- GLID: Gated Local Intrinsic Dimension Repairs the Blind Spots of Face-Forgery Detectors
- KALE: Kernel Alignment with Loss Equilibration for Stable CLIP-DINOv2 Alignment at Web Scale
- GATE-3D: Geometry-Aware Test-time Adaptive Reranking for Open-Set 3D Shape Retrieval
- Approximate Nearest Neighbor Search for Modern AI: A Projection-Augmented Graph Approach
- The JEPA Predictor: A Transferable Operator for Occluded Feature Completion
- Appearance Pointers -- Multimodal Region Control of Diffusion Transformers
- GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling
- SemDINO: DINOv3-Guided Cross-Temporal Semantic Alignment Network for Remote Sensing Change Detection
- ShotPlan: Cinematic Video Generation with Learnable Planning Token
- Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation Detection
- VGOcc: Learning Visual-Geometric Gaussians for Vision-Centric 3D Driving Occupancy Prediction
- Memory-Supported Synergistic Adaptation for Training-Free Test-Time Medical Image Segmentation
- BrainNext: A General-Purpose Self-Supervised Foundation Model for Brain MRI Analysis
- The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric
- Patch Policy: Efficient Embodied Control via Dense Visual Representations
- EVOLVE: Efficient Learned Volume Compression with Variable-Rate Encoding on a Cross-Domain Database
- Emergent Region-Level Facial Correspondence in Frozen Vision Foundation Models
- MuViSeg: Multi-View Segment Correspondences from Dense Geometry Priors
- CANDOR: Chance-Calibrated Discordance in Frozen Foundation Encoders
- Distilling Global Traversability Priors for Image-based Affordance Prediction in Off-road Environments
- Enhancing Vision Foundation Models via Multimodal Continual Pre-Training
- Cross-Coordinate Correspondence Pruning for Image-to-Point Cloud Registration
- Induce to Empower: Improving Lightweight Baselines via Foundation Model Induction for Generalized Polyp Segmentation
- UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attenders
- MemoBench: Benchmarking World Modeling in Dynamically Changing Environments
- REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion
- Foundation-Assisted Active Learning for Object Detection Annotation
- Supervised Reward Inference
- Are All Tokens Necessary for Visual Place Recognition? An Empirical Study of Token Reduction for Efficient Inference
- The Role of Initialization in 3D Gaussian Splatting
- DPNeXt: A Lightweight Multi-Scale Feature Fusion Framework for Efficient ViT-Based Multi-Task Dense Prediction
- Monkey King Bang: A Unified Scientific Multimodal Foundation Model
- RhinoVLA Technical Report
- Orbis 2: A Hierarchical World Model for Driving
- NoDrift3R: Raymap-Guided Coupling for Drift-Robust Unposed Feed-Forward 3D Reconstruction
- From Reconstruction to Interpretation: Zero-Setup Multi-Phase Segmentation of X-ray Tomography Data
- DS@GT ARC at AnimalCLEF 2026: Species-Aware Graph Construction for Multi-Species Animal Re-Identification
- Geometry-Enhanced Portion Estimation for Multimodal LLMs
- Do Vision Encoders Exhibit Human-like Color Thresholds?
- Distributional Matching for Vector Quantization: A Unified Theoretical and Empirical Framework
- Exo2EgoPose: Leveraging Exocentric Demonstrations for Vision-Language guided Egocentric 3D Hand Pose Forecasting
- Pretraining Multiple Instance Learning Networks with Multi-Teacher Distillation from Pathology Slide Foundation Models
- NavCMPO: Critic-Guided MeanFlow Policy Optimization for Adaptive Navigation
- DINO-SLAM: DINO-informed RGB-D SLAM for Neural Implicit and Explicit Representations
- Training-Free Metrics for Synthetic Object Detection Data: A Proxy for Detector Performance
- DriftWorld: Fast World Modeling through Drifting
- CoDi -- an exemplar-conditioned diffusion model for low-shot counting
- Symbal: Detecting Systematic Misalignments in Model-Generated Captions
- Multimodality as Supervision: Self-Supervised Specialization to the Test Environment via Multimodality
- Implicit Virtual Leader: Decentralized Vision-Only Relative Pose Estimation for Multi-Robot Formations
- The Third Competition on Document Forgery Detection on ID-Cards and Passports
- MIDI-RAE-JEPA: Hierarchical Representation Learning and Generation for Symbolic Music
- GlobalForge: Towards Robust AI-Generated Image Detection
- FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with Transformers
- Multi-Scale ViT Inference with Habitat-Fit Priors and kNN Retrieval for Multi-Species Plant Identification
- Parameter-efficient Prompt Tuning of Vision Foundation Model With Adaptive Focal Loss for Interpretable MCI Screening
- One-for-All Adaptive Radiotherapy Planning Agent: A Foundation Framework for Daily CBCT-guided Radiotherapy
- Selectivity Drives Efficiency: Dataset Pruning for Visual Place Recognition
- EmbodiedDiffusion: End-to-End Traversability-Guided Visual Diffusion for Heterogeneous Robot Navigation
- On the Role of Depth in Surgical Vision Foundation Models: An Empirical Study of RGB-D Pre-training
- Causal Supervision of Attention for Affective Behaviour Analysis
- Towards Consistent Video Geometry Estimation
- FISHER: Gradient-Decoupled Hierarchical Multi-Task Learning for Fine-Grained Aquatic Species Recognition
- Billions of Sketches Reveal Hidden Cultural Variation in Human Concepts
- Instant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene Simulation
- Flow Matching in Feature Space for Stochastic World Modeling
- PoseIDON: 6DoF pose estimation with foundation model features for marine sediment burial mapping
- FedCC: A Low-Resource Federated Adaptation of Foundation Models for Robust Corpus Callosum localization in Fetal Ultrasound Images
- AlphaOracle: Oracle bone script decipherment via human-workflow-inspired deep learning
- What Makes Linguistic Representations Good Models of High-Level Visual Perception in the Human Brain?
- Harnessing citizen science to contextualize adaptation mechanism discovery
- TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward
- Why AI systems don't learn and what to do about it: Lessons on autonomous learning from cognitive science
- One View Is Enough! Monocular Training for In-the-Wild Novel View Generation
- Cross-Modal Taxonomic Generalization in (Vision-) Language Models
- Human-level 3D shape perception emerges from multi-view learning
- Vision-Language Model Purified Semi-Supervised Semantic Segmentation for Remote Sensing Images
- STAR: Skeletal Token Alignment and Rearrangement for Interaction Recognition
- MuM: Multi-View Masked Image Modeling for 3D Vision
- Speedrunning ImageNet Diffusion
- MRD: Using Physically Based Differentiable Rendering to Probe Vision Models for 3D Scene Understanding
- Predicting upcoming visual features during eye movements yields scene representations aligned with human visual cortex
- Depth Anything 3: Recovering the Visual Space from Any Views
- What We Don't C: Manifold Disentanglement for Structured Discovery
- Chimère Ω — blueprint for a physico-cognitively inspired local-first LLM runtime
- SAM 3D: 3Dfy Anything in Images
- Human Mesh Modeling for Anny Body
- Kinaema: a recurrent sequence model for memory and pose in motion
- PointSt3R: Point Tracking through 3D Grounded Correspondence
- Words That Make Language Models Perceive
- Improving Robotic Manipulation with Efficient Geometry-Aware Vision Encoder
- Large Vision Models Can Solve Mental Rotation Problems
- MapAnything: Universal Feed-Forward Metric 3D Reconstruction
- Toward comprehensive cellular characterization of H&E slides
- Alligat0R: Pre-Training Through Co-Visibility Segmentation for Relative Camera Pose Regression
- Towards Early Detection: AI-Based Five-Year Forecasting of Breast Cancer Risk Using Digital Breast Tomosynthesis Imaging
- What does really matter in image goal navigation?
- Token Bottleneck: One Token to Remember Dynamics
- PanSt3R: Multi-view Consistent Panoptic Segmentation
- QuARI: Query Adaptive Retrieval Improvement
- On the Closed-Form of Flow Matching: Generalization Does Not Arise from Target Stochasticity
- Quick ViTs: Speeding up Vision Transformers through Equivariance
- Perception Encoder: The best visual embeddings are not at the output of the network
- Visual Language Models show widespread visual deficits on neuropsychological tests
- OmniSVG: A Unified Scalable Vector Graphics Generation Model
- Entropic Time Schedulers for Generative Diffusion Models
- RANa: Retrieval-Augmented Navigation
- BOP Challenge 2024 on Model-Based and Model-Free 6D Object Pose Estimation
- VGGT: Visual Geometry Grounded Transformer
- DUNE: Distilling a Universal Encoder from Heterogeneous 2D and 3D Teachers
- From Pixels to Components: Eigenvector Masking for Visual Representation Learning
- How far can we go with ImageNet for Text-to-Image generation?
- ILIAS: Instance-Level Image retrieval At Scale
- Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance
- Rethinking the Use of Vision Transformers for AI-Generated Image Detection
- Omni-Attribute: Open-vocabulary Attribute Encoder for Visual Concept Personalization
- Distribution Matching Variational AutoEncoder
- ELViS: Efficient Visual Similarity from Local Descriptors that Generalizes Across Domains
- Next Best View Selections for Semantic and Dynamic 3D Gaussian Splatting
- Memorization in 3D Shape Generation: An Empirical Study
- VL-RouterBench: A Benchmark for Vision-Language Model Routing
- Targeted Linguistic Analysis of Sign Language Models with Minimal Translation Pairs
- A Unified Framework for Vision Transformers Equivariant to Discrete Subgroups of O(2)
- VGGT-Ω
- Video Understanding: From Geometry and Semantics to Unified Models
- BBoxMaskPose v2: Expanding Mutual Conditioning to 3D
- RS-Prune: Training-Free Data Pruning at High Ratios for Efficient Remote Sensing Diffusion Foundation Models
- GVSynergy-Det: Synergistic Gaussian-Voxel Representations for Multi-View 3D Object Detection
- 3D sans 3D Scans: Scalable Pre-training from Video-Generated Point Clouds
- Toward Stable Semi-Supervised Remote Sensing Segmentation via Co-Guidance and Co-Fusion
- Revisiting [CLS] and Patch Token Interaction in Vision Transformers
- Split4D: Decomposed 4D Scene Reconstruction Without Video Segmentation
- Retrieve and Segment: Are a Few Examples Enough to Bridge the Supervision Gap in Open-Vocabulary Segmentation?
- ShapeR: Robust Conditional 3D Shape Generation from Casual Captures
- Vision Transformers Need Registers
- EgoReAct: Egocentric Video-Driven 3D Human Reaction Generation
- Multimodal Diffeomorphic Registration with Neural ODEs and Structural Descriptors
- Visual Autoregressive Modelling for Monocular Depth Estimation
- Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- Open-Source Multimodal Moxin Models with Moxin-VLM and Moxin-VLA
- Manifold-Constrained Noise Optimization for Diverse Diffusion Sampling
- Multimodal Semantic-Probabilistic Objectness for Open World Object Detection
- DuoAD: Leveraging [CLS] Dual Characteristics for Training-Free Few-Shot Anomaly Detection
- XMIX: Combating Extremely Noisy Labels via Local Smoothness in Self-Supervised Feature Space
- Unlocking the Power of Critical Factors for 3D Visual Geometry Estimation
- Are Vision Foundation Models Foundational for Electron Microscopy Image Segmentation?
- Compressing Observation History into Agent Memory: Distilling Transformers into Recurrent Transformers
- Being-H0.7: A Latent World-Action Model from Egocentric Videos
- GaitFace: A Multimodal Dataset for Long-Range Person Identification
- Pixal3D: Pixel-Aligned 3D Generation from Images
- Urban visual uniqueness: A landmark-free framework to quantify city's identity and distinctiveness from everyday scenes
- Robust 6-DoF Object Pose Tracking with Built-In Recovery under Occlusions and Rapid Object Motions
- DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification
- N0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation
- N0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens
- The potential for AI to revolutionize conservation: a horizon scan
- RRTrack: Robust and Recoverable Object 6D Pose Tracking for Dynamic Scenes
- UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models
- KAI: A Kinematic-Aware Interface for Data-Efficient Articulated Object Manipulation
- IJCB-AFMFR 2026: Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data
- MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning
- ESRVS: Extreme Semi-Supervised Retinal Vessel Segmentation with a Single Annotated Image
- Orient Anything V2: Unifying Orientation and Rotation Understanding
- A Scale-adaptive Vision Model Links C. elegans Neuronal Morphology to Behavior for Neurotoxicity Assessment
- Fashion-3DLR: A Controllable 3D Garment Generation Using Pairwise Fashion Elements for Intelligent Design
- ScaleResfusion: Residual Rectified Flow based on Residual Vector Field
- DisasterTD: Disaster Toponym Disambiguation Using Multimodal LLMs and Cross-View Geolocalization
- ReLATE: Reliability-Guided Evidence Fusion for Robust UAV--Satellite cross-view Geo-Localization
- Metric Surface Reconstruction of Neurosurgical Scenes from Monocular Operating Microscope Images and Microscope Pose
- Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
- iSHIFT: Lightweight Slow-Fast GUI Agent with Adaptive Perception
- Robustifying pathology foundation models via fine-tuning
- JPEG AIC2026: A large-scale dataset for fine-grained assessment of image coding
- Controlling Embedding Spaces with Text-Conditioned Transformations
- Gaze-Anchored Social Net: Decoding Implicit Relations via Joint Modeling
- Cortex: Compact Behavior Cloning for Quake with Frozen Visual Features
- Benchmarking the Domain Gap: Model Selection Instability Under Domain Shift in Video Capsule Endoscopy
- T2LDM++: A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene Generation
- Ground4D: Consistency-Aware 4D Reconstruction from Monocular Video
- TSA: Temporal Slot Activation for Persistent Object-Centric Video Representation
- ChangeFlow -- Latent Rectified Flow for Change Detection in Remote Sensing
- FREPix: Frequency-Heterogeneous Flow Matching for Pixel-Space Image Generation
- Personalize Your Large Vision-language Models With In-context Prompt Tuning
- A Cascaded Edge-Cloud Architecture for Automated Diabetic Retinopathy Screening
- When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models
- Context Sensitivity Improves Human-Machine Visual Alignment
- DreamCAD: Scaling Multi-modal CAD Generation using Differentiable Parametric Surfaces
- Self-Distillation of Hidden Layers for Self-Supervised Representation Learning
- Reverso: Efficient Time Series Foundation Models for Zero-shot Forecasting
- PixelGen: Improving Pixel Diffusion with Perceptual Supervision
- InstructMoLE: Instruction-Guided Mixture of Low-rank Experts for Multi-Conditional Image Generation
- StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
- Reloc-VGGT: Visual Re-localization with Geometry Grounded Transformer
- Pro-Pose: Unpaired Full-Body Portrait Synthesis via Canonical UV Maps
- ImagineNav++: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination
- LogicLens: Visual-Logical Co-Reasoning for Text-Centric Forgery Analysis
- RLLaVA: An RL-central Framework for Language and Vision Assistants
- TICON: A Slide-Level Tile Contextualizer for Histopathology Representation Learning
- Surgical Scene Segmentation using a Spike-Driven Video Transformer with Real-Time Potential
- UltraShape 1.0: High-Fidelity 3D Shape Generation via Scalable Geometric Refinement
- Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding
- Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
- MVInverse: Feed-forward Multi-view Inverse Rendering in Seconds
- PanoGrounder: Bridging 2D and 3D with Panoramic Scene Representations for VLM-based 3D Visual Grounding
- UniPR-3D: Towards Universal Visual Place Recognition with Visual Geometry Grounded Transformer
- CHAMMI-75: Pre-training multi-channel models with heterogeneous microscopy images
- Towards Generative Location Awareness for Disaster Response: A Probabilistic Cross-view Geolocalization Approach
- AlignPose: Generalizable 6D Pose Estimation via Multi-view Feature-metric Alignment
- UTDesign: A Unified Framework for Stylized Text Editing and Generation in Graphic Design Images
- Evolutionary Neural Architecture Search with Dual Contrastive Learning
- Detecting Non-Optimal Decisions of Embodied Agents via Diversity-Guided Metamorphic Testing
- PairFlow: Closed-Form Source-Target Coupling for Few-Step Generation in Discrete Flow Models
- Bring My Cup! Personalizing Vision-Language-Action Models with Visual Attentive Prompting
- WSD-MIL: Window Scale Decay Multiple Instance Learning for Whole Slide Image Classification
- Zero-Shot Segmentation through Prototype-Guidance for Multi-Label Plant Species Identification
- How Much 3D Do Video Foundation Models Encode?
- SE360: Semantic Edit in 360^∘ Panoramas via Hierarchical Data Construction
- Block-Recurrent Dynamics in Vision Transformers
- Probing and Leveraging Video Diffusion Transformer Features for Robust Point Tracking
- Pix2NPHM: Learning to Regress NPHM Reconstructions From a Single Image
- SemanticGen: Video Generation in Semantic Space
- The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding
- LoGoPlanner: Localization Grounded Navigation Policy with Metric-aware Visual Geometry
- A Flexible Field-Based Policy Learning Framework for Diverse Robotic Systems and Sensors
- WaTeRFlow: Watermark Temporal Robustness via Flow Consistency
- VLNVerse: A Benchmark for Vision-Language Navigation with Versatile, Embodied, Realistic Simulation and Evaluation
- Finer-Personalization Rank: Fine-Grained Retrieval Examines Identity Preservation for Personalized Generation
- MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture
- 3SGen: Unified Subject, Style, and Structure-Driven Image Generation with Adaptive Task-specific Memory
- RMLer: Synthesizing Novel Objects across Diverse Categories via Reinforcement Mixing Learning
- Is Visual Realism Enough? Evaluating Gait Biometric Fidelity in Generative AI Human Animation
- Eff-GRot: Efficient and Generalizable Rotation Estimation with Transformers
- SmartSight: Mitigating Hallucination in Video-LLMs Without Compromising Video Understanding via Temporal Attention Collapse
- WoundNet-Ensemble: A Novel IoMT System Integrating Self-Supervised Deep Learning and Multi-Model Fusion for Automated, High-Accuracy Wound Classification and Healing Progression Monitoring
- UniMPR: A Unified Framework for Multimodal Place Recognition with Heterogeneous Sensor Configurations
- SG-RIFE: Semantic-Guided Real-Time Intermediate Flow Estimation with Diffusion-Competitive Perceptual Quality
- Atlas is Your Perfect Context: One-Shot Customization for Generalizable Foundational Medical Image Segmentation
- Uncertainty-Gated Region-Level Retrieval for Robust Semantic Segmentation
- ReDepth Anything: Test-Time Depth Refinement via Self-Supervised Re-lighting
- Keypoint Counting Classifiers: Turning Vision Transformers into Self-Explainable Models Without Training
- ClothHMR: 3D Mesh Recovery of Humans in Diverse Clothing from Single Image
- MMLANDMARKS: a Cross-View Instance-Level Benchmark for Geo-Spatial Understanding
- InsertAnywhere: Geometrically Grounded and Optics-Aware Video Object Insertion
- Robust Scene Coordinate Regression via Geometrically-Consistent Global Descriptors
- FLEG: Feed-Forward Language Embedded Gaussian Splatting from Any Views via Compact Semantic Representation
- RoomEditor++: A Parameter-Sharing Diffusion Architecture for High-Fidelity Furniture Synthesis
- Next-Embedding Prediction Makes Strong Vision Learners
- SFTok: Bridging the Performance Gap in Discrete Tokenizers
- SceneDiff: A Benchmark and Method for Multiview Object Change Detection
- Instant Expressive Gaussian Head Avatars at Over 100 FPS
- Causal-Tune: Mining Causal Factors from Vision Foundation Models for Domain Generalized Semantic Segmentation
- Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models
- Sceniris: A Fast Procedural Scene Generation Framework
- SegGraph: Leveraging Graphs of SAM Segments for Few-Shot 3D Part Segmentation
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- In Pursuit of Pixel Supervision for Visual Pre-training
- FlexAvatar: Learning Complete 3D Head Avatars with Partial Supervision
- MoonSeg3R: Monocular Online Zero-Shot Segment Anything in 3D with Reconstructive Foundation Priors
- MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training
- The Deleuzian Representation Hypothesis
- Null-LoRA: Low-Rank Adaptation on Null Space
- BEV-Patch-PF: Particle Filtering with BEV-Aerial Feature Matching for Off-Road Geo-Localization
- Off The Grid: Detection of Primitives for Feed-Forward 3D Gaussian Splatting
- Artificial Intelligence for the Assessment of Peritoneal Carcinosis during Diagnostic Laparoscopy for Advanced Ovarian Cancer
- Magnification-Aware Distillation (MAD): A Self-Supervised Framework for Unified Representation Learning in Gigapixel Whole-Slide Images
- PANDA-PLUS-Bench: A Clinical Benchmark for Evaluating Robustness of AI Foundation Models in Prostate Cancer Diagnosis
- Spherical Leech Quantization for Visual Tokenization and Generation
- ART: Articulated Reconstruction Transformer
- Semantic search for 100M+ galaxy images using AI-generated captions
- SS4D: Native 4D Generative Model via Structured Spacetime Latents
- Towards Test-time Efficient Visual Place Recognition via Asymmetric Query Processing
- SketchAssist: A Practical Assistant for Semantic Edits and Precise Local Redrawing
- Consistent Instance Field for Dynamic Scene Understanding
- EXAONE Path 2.5: Pathology Foundation Model with Multi-Omics Alignment
- LCMem: A Universal Model for Robust Image Memorization Detection
- SuperCLIP: CLIP with Simple Classification Supervision
- Native Intelligence Emerges from Large-Scale Clinical Practice: A Retinal Foundation Model with Deployment Efficiency
- EchoVLM: Measurement-Grounded Multimodal Learning for Echocardiography
- Recurrent Video Masked Autoencoders
- Directional Textual Inversion for Personalized Text-to-Image Generation
- World Models Can Leverage Human Videos for Dexterous Manipulation
- Image Diffusion Preview with Consistency Solver
- Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation
- RecTok: Reconstruction Distillation along Rectified Flow
- DePT3R: Joint Dense Point Tracking and 3D Reconstruction of Dynamic Scenes in a Single Forward Pass
- Forging a Dynamic Memory: Retrieval-Guided Continual Learning for Generalist Medical Foundation Models
- Revisiting 2D Foundation Models for Scalable 3D Medical Image Classification
- DBT-DINO
- Unlocking Generalization in Polyp Segmentation with DINO Self-Attention "keys"
- Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling
- Patch-wise Retrieval: A Bag of Practical Techniques for Instance-level Matching
- RePack then Refine: Efficient Diffusion Transformer with Vision Foundation Model
- SpaceControl: Introducing Test-Time Spatial Control to 3D Generative Modeling
- Brain-Semantoks: Learning Semantic Tokens of Brain Dynamics with a Self-Distilled Foundation Model
- Reconstruction as a Bridge for Event-Based Visual Question Answering
- On Geometric Understanding and Learned Priors in Feed-forward 3D Reconstruction Models
- JoyAvatar: Real-time and Infinite Audio-Driven Avatar Generation with Autoregressive Diffusion
- Collaborative Reconstruction and Repair for Multi-class Industrial Anomaly Detection
- VFMF: World Modeling by Forecasting Vision Foundation Model Features
- An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
- AutoRefiner: Improving Autoregressive Video Diffusion Models via Reflective Refinement Over the Stochastic Sampling Path
- AlcheMinT: Fine-grained Temporal Control for Multi-Reference Consistent Video Generation
- Fast-FoundationStereo: Real-Time Zero-Shot Stereo Matching
- Learning Category-level Last-meter Navigation from RGB Demonstrations of a Single-instance
- WholeBodyVLA: Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control
- SceneMaker: Open-set 3D Scene Generation with Decoupled De-occlusion and Pose Estimation Model
- Empowering Dynamic Urban Navigation with Stereo and Mid-Level Vision
- Group Diffusion: Enhancing Image Generation by Unlocking Cross-Sample Collaboration
- Bidirectional Normalizing Flow: From Data to Noise and Back
- E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-training
- ClusIR: Towards Cluster-Guided All-in-One Image Restoration
- Towards Efficient and Effective Multi-Camera Encoding for End-to-End Driving
- Any4D: Unified Feed-Forward Metric 4D Reconstruction
- BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models
- DuetSVG: Unified Multimodal SVG Generation with Internal Visual Guidance
- SoccerMaster: A Vision Foundation Model for Soccer Understanding
- Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
- What matters for Representation Alignment: Global Information or Spatial Structure?
- Self-Supervised Contrastive Embedding Adaptation for Endoscopic Image Matching
- Geo6DPose: Fast Zero-Shot 6D Object Pose Estimation via Geometry-Filtered Feature Matching
- ShotDirector: Directorially Controllable Multi-Shot Video Generation with Cinematographic Transitions
- THE-Pose: Topological Prior with Hybrid Graph Fusion for Estimating Category-Level 6D Object Pose
- Solving Semi-Supervised Few-Shot Learning from an Auto-Annotation Perspective
- Latent Chain-of-Thought World Modeling for End-to-End Driving
- MoCapAnything: Unified 3D Motion Capture for Arbitrary Skeletons from Monocular Videos
- Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors
- DynaIP: Dynamic Image Prompt Adapter for Scalable Zero-shot Personalized Text-to-Image Generation
- HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models
- GAINS: Gaussian-based Inverse Rendering from Sparse Multi-View Captures
- Simultaneous Tactile-Visual Perception for Learning Multimodal Robot Manipulation
- GLaD: Geometric Latent Distillation for Vision-Language-Action Models
- UrbanNav: Learning Language-Guided Urban Navigation from Web-Scale Human Trajectories
- Cytoplasmic Strings Analysis in Human Embryo Time-Lapse Videos using Deep Learning Framework
- UniPart: Part-Level 3D Generation with Unified 3D Geom-Seg Latents
- Geometry-to-Image Synthesis-Driven Generative Point Cloud Registration
- From SAM to DINOv2: Towards Distilling Foundation Models to Lightweight Baselines for Generalized Polyp Segmentation
- CytoDINO: Risk-Aware and Biologically-Informed Adaptation of DINOv3 for Bone Marrow Cytomorphology
- Can TabPFN Compete with GNNs for Node Classification via Graph Tabularization?
- A Multi-Robot Platform for Robotic Triage Combining Onboard Sensing and Foundation Models
- SegEarth-OV3: Exploring SAM 3 for Open-Vocabulary Semantic Segmentation in Remote Sensing Images
- OpenMonoGS-SLAM: Monocular Gaussian Splatting SLAM with Open-set Semantics
- Inferring Compositional 4D Scenes without Ever Seeing One
- OpenSubject: Leveraging Video-Derived Identity and Diversity Priors for Subject-driven Image Generation and Manipulation
- Accuracy Does Not Guarantee Human-Likeness in Monocular Depth Estimators
- Repulsor: Accelerating Generative Modeling with a Contrastive Memory Bank
- ConceptPose: Training-Free Zero-Shot Object Pose Estimation using Concept Vectors
- OSMO: Open-Source Tactile Glove for Human-to-Robot Skill Transfer
- Trajectory Densification and Depth from Perspective-based Blur
- Selfi: Self Improving Reconstruction Engine via 3D Geometric Feature Alignment
- Is Generation Required for Data-Efficient Perception?
- OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory
- GorillaWatch: An Automated System for In-the-Wild Gorilla Re-Identification and Population Monitoring
- ViSA: 3D-Aware Video Shading for Real-Time Upper-Body Avatar Creation
- A Geometric Unification of Concept Learning with Concept Cones
- VFM-VLM: Vision Foundation Model and Vision Language Model based Visual Comparison for 3D Pose Estimation
- D3-Predictor: Noise-Free Deterministic Diffusion for Dense Prediction
- Online Segment Any 3D Thing as Instance Tracking
- See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations
- Reconstructing Objects along Hand Interaction Timelines in Egocentric Video
- From Orbit to Ground: Generative City Photogrammetry from Extreme Off-Nadir Satellite Images
- RVLF: A Reinforcing Vision-Language Framework for Gloss-Free Sign Language Translation
- Towards Robust Pseudo-Label Learning in Semantic Segmentation: An Encoding Perspective
- Dynamic Visual SLAM using a General 3D Prior
- Generalized Geometry Encoding Volume for Real-time Stereo Matching
- VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
- AGORA: Adversarial Generation Of Real-time Animatable 3D Gaussian Head Avatars
- SpectraIrisPAD: Leveraging Vision Foundation Models for Spectrally Conditioned Multispectral Iris Presentation Attack Detection
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
- ReCAD: Reinforcement Learning Enhanced Parametric CAD Model Generation with Vision-Language Models
- Learning High-Fidelity Cloth Animation via Skinning-Free Image Transfer
- MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models
- See in Depth: Training-Free Surgical Scene Segmentation with Monocular Depth Priors
- Rethinking Infrared Small Target Detection: A Foundation-Driven Efficient Paradigm
- EmoStyle: Emotion-Driven Image Stylization
- LoC-Path: Learning to Compress for Pathology Multimodal Large Language Models
- Group Orthogonal Low-Rank Adaptation for RGB-T Tracking
- Object Reconstruction under Occlusion with Generative Priors and Contact-induced Constraints
- 4DLangVGGT: 4D Language-Visual Geometry Grounded Transformer
- Stable Single-Pixel Contrastive Learning for Semantic and Geometric Tasks
- LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token Merging
- Semantics Lead the Way: Harmonizing Semantic and Texture Modeling with Asynchronous Latent Diffusion
- Paris 2.0: A Decentralized Diffusion Model for Video Generation
- EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation
- LaFiTe: A Generative Latent Field for 3D Native Texturing
- Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length
- Denoise to Track: Harnessing Video Diffusion Priors for Robust Correspondence
- Refaçade: Editing Object with Given Reference Texture
- Not All Birds Look The Same: Identity-Preserving Generation For Birds
- DeRA: Decoupled Representation Alignment for Video Tokenization
- SEASON: Mitigating Temporal Hallucination in Video Large Language Models via Self-Diagnostic Contrastive Decoding
- Look Around and Pay Attention: Multi-camera Point Tracking Reimagined with Transformers
- Unique Lives, Shared World: Learning from Single-Life Videos
- Domain Feature Collapse: Implications for Out-of-Distribution Detection and Solutions
- C3G: Learning Compact 3D Representations with 2K Gaussians
- PosA-VLA: Enhancing Action Generation via Pose-Conditioned Anchor Attention
- VAT: Vision Action Transformer by Unlocking Full Representation of ViT
- AaPE: Aliasing-aware Patch Embedding for Self-Supervised Audio Representation Learning
- AdaPower: Specializing World Foundation Models for Predictive Manipulation
- OpenTrack3D: Towards Accurate and Generalizable Open-Vocabulary 3D Instance Segmentation
- CSMapping: Scalable Crowdsourced Semantic Mapping and Topology Inference for Autonomous Driving
- Exploiting Domain Properties in Language-Driven Domain Generalization for Semantic Segmentation
- What Is The Best 3D Scene Representation for Robotics? From Geometric to Foundation Models
- ShelfGaussian: Shelf-Supervised Open-Vocabulary Gaussian-based 3D Scene Understanding
- Network of Theseus (like the ship)
- Emergent Outlier View Rejection in Visual Geometry Grounded Transformers
- Vision Foundry: A System for Training Foundational Vision AI Models
- Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding
- Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?
- Flux4D: Flow-based Unsupervised 4D Reconstruction
- CAMEO: Correspondence-Attention Alignment for Multi-View Diffusion Models
- Video2Act: A Dual-System Video Diffusion Policy with Robotic Spatio-Motional Modeling
- MultiShotMaster: A Controllable Multi-Shot Video Generation Framework
- Hierarchical Process Reward Models are Symbolic Vision Learners
- TEXTRIX: Latent Attribute Grid for Native Texture Generation and Beyond
- Polar Perspectives: Evaluating 2-D LiDAR Projections for Robust Place Recognition with Visual Foundation Models
- From Navigation to Refinement: Revealing the Two-Stage Nature of Flow-based Diffusion Models through Oracle Velocity
- FiMMIA: scaling semantic perturbation-based membership inference across modalities
- Unsupervised Structural Scene Decomposition via Foreground-Aware Slot Attention with Pseudo-Mask Guidance
- AVGGT: Rethinking Global Attention for Accelerating VGGT
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- KM-ViPE: Online Tightly Coupled Vision-Language-Geometry Fusion for Open-Vocabulary Semantic SLAM
- Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos
- FreqEdit: Preserving High-Frequency Features for Robust Multi-Turn Image Editing
- SPARK: Sim-ready Part-level Articulated Reconstruction with VLM Knowledge
- No Trust Issues Here: A Technical Report on the Winning Solutions for the Rayan AI Contest
- OpenBox: Annotate Any Bounding Boxes in 3D
- PSR: Scaling Multi-Subject Personalized Image Generation with Pairwise Subject-Consistency Rewards
- TabletopGen: Tabletop Scene Generation and Interactive Simulation for Robotic Manipulation
- Assimilation Matters: Model-level Backdoor Detection in Vision-Language Pretrained Models
- Real-World Reinforcement Learning of Active Perception Behaviors
- IGen: Scalable Data Generation for Robot Learning from Open-World Images
- Fantastic Features and Where to Find Them: A Probing Method to combine Features from Multiple Foundation Models
- Learning Eigenstructures of Unstructured Data Manifolds
- PAGen: Phase-guided Amplitude Generation for Domain-adaptive Object Detection
- FOM-Nav: Frontier-Object Maps for Object Goal Navigation
- PhotoFramer: Multi-modal Image Composition Instruction
- TAP-CT: 3D Task-Agnostic Pretraining of Computed Tomography Foundation Models
- Dynamic-eDiTor: Training-Free Text-Driven 4D Scene Editing with Multimodal Diffusion Transformer
- VFM-ISRefiner: Towards Better Adapting Vision Foundation Models for Interactive Segmentation of Remote Sensing Images
- Silhouette-based Gait Foundation Model
- Cosine-Similarity Methods for Efficient Training and Sampling in High-Dimensional Latent Spaces
- Describe Anything Anywhere At Any Moment
- Asset-Driven Sematic Reconstruction of Dynamic Scene with Multi-Human-Object Interactions
- PhysGen: Physically Grounded 3D Shape Generation for Industrial Design
- DisMo: Disentangled Motion Representations for Open-World Motion Transfer
- GENA3D: Generative Amodal 3D Modeling by Bridging 2D Priors and 3D Coherence
- DEAL-300K: Diffusion-based Editing Area Localization with a 300K-Scale Dataset and Frequency-Prompted Baseline
- Pathryoshka: Compressing Pathology Foundation Models via Multi-Teacher Knowledge Distillation with Nested Embeddings
- VG3T: Visual Geometry Grounded Gaussian Transformer
- GSPN-2: Efficient Parallel Sequence Modeling
- Overcoming the Curvature Bottleneck in MeanFlow
- Learning to Predict Aboveground Biomass from RGB Images with 3D Synthetic Scenes
- Contrastive Heliophysical Image Pretraining for Solar Dynamics Observatory Records
- Continual Error Correction on Low-Resource Devices
- SHIC-XE: Viewpoint-Invariant Explainability via Dense 2D-3D Correspondences: an Application to Equine Pain Recognition
- Combining Projected Uncertainty for Self-Supervised Visual Odometry: From Two-Frame to Multi-Frame
- Adversarial Flow Models
- Rethinking Cross-Generator Image Forgery Detection through DINOv3
- Canvas-to-Image: Compositional Image Generation with Multimodal Controls
- Wukong's 72 Transformations: High-fidelity Textured 3D Morphing via Flow Models
- Cleaning the Pool: Progressive Filtering of Unlabeled Pools in Deep Active Learning
- Attention-Guided Patch-Wise Sparse Adversarial Attacks on Vision-Language-Action Models
- E0: Enhancing Generalization and Fine-Grained Control in VLA Models via Continuized Discrete Diffusion
- Generalized Design Choices for Deepfake Detectors
- Uni-Hema: Unified Model for Digital Hematopathology
- MeanFlow Transformers with Representation Autoencoders
- Multi-modal On-Device Learning for Monocular Depth Estimation on Ultra-low-power MCUs
- A Probabilistic Framework for Temporal Distribution Generalization in Industry-Scale Recommender Systems
- When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action Models
- DINO-Tok: Adapting DINO for Visual Tokenizers
- PixelDiT: Pixel Diffusion Transformers for Image Generation
- Image2Gcode: Image-to-G-code Generation for Additive Manufacturing Using Diffusion-Transformer Model
- Automated Histopathologic Assessment of Hirschsprung Disease Using a Multi-Stage Vision Transformer Framework
- MajutsuCity: Language-driven Aesthetic-adaptive City Generation with Controllable 3D Assets and Layouts
- CrossEarth-Gate: Fisher-Guided Adaptive Tuning Engine for Efficient Adaptation of Cross-Domain Remote Sensing Semantic Segmentation
- The Image as Its Own Reward: Reinforcement Learning with Adversarial Reward for Image Generation
- Patch-Level Glioblastoma Subregion Classification with a Contrastive Learning-Based Encoder
- ADNet: A Large-Scale and Extensible Multi-Domain Benchmark for Anomaly Detection Across 380 Real-World Categories
- PRADA: Probability-Ratio-Based Attribution and Detection of Autoregressive-Generated Images
- DinoLizer: Separating VAE and Diffusion Artifacts in Generative Inpainting Localization
- Cross-Contrastive Clustering for Multimodal Attributed Graphs with Dual Graph Filtering
- OmniRefiner: Reinforcement-Guided Local Diffusion Refinement
- HiCoGen: Hierarchical Compositional Text-to-Image Generation in Diffusion Models via Reinforcement Learning
- Supervise Less, See More: Training-free Nuclear Instance Segmentation with Prototype-Guided Prompting
- LumiTex: Towards High-Fidelity PBR Texture Generation with Illumination Context
- Unifying Perception and Action: A Hybrid-Modality Pipeline with Implicit Visual Chain-of-Thought for Robotic Action Generation
- ShapeGen: Towards High-Quality 3D Shape Synthesis
- Vision-Language Enhanced Foundation Model for Semi-supervised Medical Image Segmentation
- SkillSight: Efficient First-Person Skill Assessment with Gaze
- Learning Massively Multitask World Models for Continuous Control
- UniGame: Turning a Unified Multimodal Model Into Its Own Adversary
- DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation
- A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene Generation
- View-Consistent Diffusion Representations for 3D-Consistent Video Generation
- Eevee: Towards Close-up High-resolution Video-based Virtual Try-on
- Learning Scalable Temporal Representations in Spiking Neural Networks Without Labels
- Perceptual Taxonomy: Evaluating and Guiding Hierarchical Scene Reasoning in Vision-Language Models
- DetAny4D: Detect Anything 4D Temporally in a Streaming RGB Video
- Rethinking Garment Conditioning in Diffusion-based Virtual Try-On
- Seeing What Matters: Visual Preference Policy Optimization for Visual Generation
- IDEAL-M3D: Instance Diversity-Enriched Active Learning for Monocular 3D Detection
- VeCoR -- Velocity Contrastive Regularization for Flow Matching
- Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
- Diverse Instance Generation via Diffusion Models for Enhanced Few-Shot Object Detection in Remote Sensing Images
- NeAR: Coupled Neural Asset-Renderer Stack
- 4D-VGGT: A General Foundation Model with SpatioTemporal Awareness for Dynamic Scene Geometry Estimation
- RNN as Linear Transformer: A Closer Investigation into Representational Potentials of Visual Mamba Models
- ConsistCompose: Unified Multimodal Layout Control for Image Composition
- stable-pretraining-v1: Foundation Model Research Made Simple
- C3Po: Cross-View Cross-Modality Correspondence by Pointmap Prediction
- Object-centric Task Representation and Transfer using Diffused Orientation Fields
- MotionDuet: Dual-Conditioned 3D Human Motion Generation with Video-Regularized Text Learning
- Off-Road Navigation via Implicit Neural Representation of Terrain Traversability
- Observer-Actor: Active Vision Imitation Learning with Sparse-View Gaussian Splatting
- MVS-TTA: Test-Time Adaptation for Multi-View Stereo via Meta-Auxiliary Learning
- Muskie: Multi-view Masked Image Modeling for 3D Vision Pre-training
- AdaPerceiver: Transformers with Adaptive Width, Depth, and Tokens
- IE-Critic-R1: Advancing the Explanatory Measurement of Text-Driven Image Editing for Human Perception Alignment
- RAISECity: A Multimodal Agent Framework for Reality-Aligned 3D World Generation at City-Scale
- Hierarchical Semi-Supervised Active Learning for Remote Sensing
- Show Me: Unifying Instructional Image and Video Generation with Diffusion Models
- A Stitch in Time: Learning Procedural Workflow via Self-Supervised Plackett-Luce Ranking
- Attention Guided Alignment in Efficient Vision-Language Models
- SPIDER: Spatial Image CorresponDence Estimator for Robust Calibration
- Improving Multimodal Distillation for 3D Semantic Segmentation under Domain Shift
- Illustrator's Depth: Monocular Layer Index Prediction for Image Decomposition
- Scaling Self-Supervised and Cross-Modal Pretraining for Volumetric CT Transformers
- Investigating self-supervised representations for audio-visual deepfake detection
- Closing the Performance Gap Between AI and Radiologists in Chest X-Ray Reporting
- MatPedia: A Universal Generative Foundation for High-Fidelity Material Synthesis
- Align & Invert: Solving Inverse Problems with Diffusion and Flow-based Models via Representation Alignment
- METIS: Multi-Source Egocentric Training for Integrated Dexterous Vision-Language-Action Model
- SpatialGeo:Boosting Spatial Reasoning in Multimodal LLMs via Geometry-Semantics Fusion
- Loomis Painter: Reconstructing the Painting Process
- Personalized Reward Modeling for Text-to-Image Generation
- LAOF: Robust Latent Action Learning with Optical Flow Constraints
- Target Refocusing via Attention Redistribution for Open-Vocabulary Semantic Segmentation: An Explainability Perspective
- Unsupervised Image Classification with Adaptive Nearest Neighbor Selection and Cluster Ensembles
- Upsample Anything: A Simple and Hard to Beat Baseline for Feature Upsampling
- BioBench: A Blueprint to Move Beyond ImageNet for Scientific ML Benchmarks
- WALDO: Where Unseen Model-based 6D Pose Estimation Meets Occlusion
- LEGO-SLAM: Language-Embedded Gaussian Optimization SLAM
- CoordAR: One-Reference 6D Pose Estimation of Novel Objects via Autoregressive Coordinate Map Generation
- SIGMMA: Hierarchical Graph-Based Multi-Scale Multi-modal Contrastive Alignment of Histopathology Image and Spatial Transcriptome
- A Dataset and Baseline for Deep Learning-Based Visual Quality Inspection in Remanufacturing
- ShelfOcc: Native 3D Supervision beyond LiDAR for Vision-Based Occupancy Estimation
- What Your Features Reveal: Data-Efficient Black-Box Feature Inversion Attack for Split DNNs
- Quant-Trim in Practice: Improved Cross-Platform Low-Bit Deployment on Edge NPUs
- SplitFlux: Learning to Decouple Content and Style from a Single Image
- Insert In Style: A Zero-Shot Generative Framework for Harmonious Cross-Domain Object Composition
- UniSER: A Foundation Model for Unified Soft Effects Removal
- DCL-SE: Dynamic Curriculum Learning for Spatiotemporal Encoding of Brain Imaging
- RoMa v2: Harder Better Faster Denser Feature Matching
- X-WIN: Building Chest Radiograph World Model via Predictive Sensing
- FarSLIP: Discovering Effective CLIP Adaptation for Fine-Grained Remote Sensing Understanding
- HMC: Learning Heterogeneous Meta-Control for Contact-Rich Loco-Manipulation
- D-PerceptCT: Deep Perceptual Enhancement for Low-Dose CT Images
- Learning to See Through a Baby's Eyes: Early Visual Diets Enable Robust Visual Intelligence in Humans and Machines
- Online Data Curation for Object Detection via Marginal Contributions to Dataset-level Average Precision
- Semantic Context Matters: Improving Conditioning for Autoregressive Models
- Training-free Detection of AI-generated images via Cropping Robustness
- AnaCP: Toward Upper-Bound Continual Learning via Analytic Contrastive Projection
- Tissue Aware Nuclei Detection and Classification Model for Histopathology Images
- TSE-Net: Semi-supervised Monocular Height Estimation from Single Remote Sensing Images
- TripleFDS: Triple Feature Disentanglement and Synthesis for Scene Text Editing
- SkyReels-Text: Fine-Grained Font-Controllable Text Editing for Poster Design
- SOMA: Feature Gradient Enhanced Affine-Flow Matching for SAR-Optical Registration
- Semantics and Content Matter: Towards Multi-Prior Hierarchical Mamba for Image Deraining
- Distribution Matching Distillation Meets Reinforcement Learning
- Towards Metric-Aware Multi-Person Mesh Recovery by Jointly Optimizing Human Crowd in Camera Space
- DINO-MX: A Modular & Flexible Framework for Self-Supervised Learning
- PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
- CLIDD: Cross-Layer Independent Deformable Description for Efficient and Discriminative Local Feature Representation
- Appreciate the View: A Task-Aware Evaluation Framework for Novel View Synthesis
- Rank-Aware Agglomeration of Foundation Models for Immunohistochemistry Image Cell Counting
- D2-VPR: A Parameter-efficient Visual-foundation-model-based Visual Place Recognition Method via Knowledge Distillation and Deformable Aggregation
- Visible Structure Retrieval for Lightweight Image-Based Relocalisation
- VLA-R: Vision-Language Action Retrieval toward Open-World End-to-End Autonomous Driving
- Changes in Real Time: Online Scene Change Detection with Multi-View Fusion
- Constructing and Interpreting Digital Twin Representations for Visual Reasoning via Reinforcement Learning
- Data-Efficient Self-Supervised Algorithms for Fine-Grained Birdsong Analysis
- Fine-Grained DINO Tuning with Dual Supervision for Face Forgery Detection
- EgoCogNav: Cognition-aware Human Egocentric Navigation
- Sat2RealCity: Geometry-Aware and Appearance-Controllable 3D Urban Generation from Satellite Imagery
- Φeat: Physically Grounded Material Feature Representation
- DoReMi: Bridging 3D Domains via Topology-Aware Domain-Representation Mixture of Experts
- Heterogeneous Complementary Distillation
- Phantom Menace: Exploring and Enhancing the Robustness of VLA Models Against Physical Sensor Attacks
- Attentive Feature Aggregation or: How Policies Learn to Stop Worrying about Robustness and Attend to Task-Relevant Visual Cues
- Fast Data Attribution for Text-to-Image Models
- SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic Manipulation
- Diversity Over Scale: Whole-Slide Image Variety Enables H&E Foundation Model Training with Fewer Patches
- Feature Quality and Adaptability of Medical Foundation Models: A Comparative Evaluation for Radiographic Classification and Segmentation
- LoG3D: Ultra-High-Resolution 3D Shape Modeling via Local-to-Global Partitioning
- Audio-VLA: Adding Contact Audio Perception to Vision-Language-Action Model for Robotic Manipulation
- From Street to Orbit: Training-Free Cross-View Retrieval via Location Semantics and LLM Guidance
- STORM: Segment, Track, and Object Re-Localization from a Single Image
- RF-DETR: Neural Architecture Search for Real-Time Detection Transformers
- TomoGraphView: 3D Medical Image Classification with Omnidirectional Slice Representations and Graph Neural Networks
- ScaleADFG: Affordance-based Dexterous Functional Grasping via Scalable Dataset
- PriVi: Towards A General-Purpose Video Model For Primate Behavior In The Wild
- SliderEdit: Continuous Image Editing with Fine-Grained Instruction Control
- CleverBirds: A Multiple-Choice Benchmark for Fine-grained Human Knowledge Tracing
- MVSMamba: Multi-View Stereo with State Space Model
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
- Representation Learning Enables Scalable Multitask Deep Reinforcement Learning
- Empowering DINO Representations for Underwater Instance Segmentation via Aligner and Prompter
- Determinism of Randomness: Prompt-Residual Seed Shaping for Diffusion Generation
- VLMDiff: Leveraging Vision-Language Models for Multi-Class Anomaly Detection with Diffusion
- Taming Identity Consistency and Prompt Diversity in Diffusion Models via Latent Concatenation and Masked Conditional Flow Matching
- WEDepth: Efficient Adaptation of World Knowledge for Monocular Depth Estimation
- EAGLE: Episodic Appearance- and Geometry-aware Memory for Unified 2D-3D Visual Query Localization in Egocentric Vision
- Exploring the Underwater World Segmentation without Extra Training
- DI3CL: Contrastive Learning With Dynamic Instances and Contour Consistency for SAR Land-Cover Classification Foundation Model
- ViPRA: Video Prediction for Robot Actions
- Detecting Generated Images by Fitting Natural Image Distributions
- Rethinking Generative Image Pretraining: How Far Are We From Scaling Up Next-Pixel Prediction?
- Wid3R: Wide Field-of-View 3D Reconstruction via Camera Model Conditioning
- StreamDiffusionV2: A Streaming System for Dynamic and Interactive Video Generation
- YoNoSplat: You Only Need One Model for Feedforward 3D Gaussian Splatting
- ClusterMine: Robust Label-Free Visual Out-Of-Distribution Detection via Concept Mining from Text Corpora
- PlantTraitNet: An Uncertainty-Aware Multimodal Framework for Global-Scale Plant Trait Inference from Citizen Science Data
- DTTNet: Improving Video Shadow Detection via Dark-Aware Guidance and Tokenized Temporal Modeling
- Generating an Image From 1,000 Words: Enhancing Text-to-Image With Structured Captions
- SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation
- FlowFeat: Pixel-Dense Embedding of Motion Profiles
- BuildingWorld: A Structured 3D Building Dataset for Urban Foundation Models
- SAMora: Enhancing SAM through Hierarchical Self-Supervised Pre-Training for Medical Images
- Affordance-Guided Coarse-to-Fine Exploration for Base Placement in Open-Vocabulary Mobile Manipulation
- Temporal-Guided Visual Foundation Models for Event-Based Vision
- Identity Card Presentation Attack Detection: A Systematic Review
- Towards Implicit Aggregation: Robust Image Representation for Place Recognition in the Transformer Era
- Adapted Foundation Models for Breast MRI Triaging in Contrast-Enhanced and Non-Contrast Enhanced Protocols
- Vision Foundation Models in Agriculture: Toward Domain-Specific Adaptation for Weed Herbicide Trials Assessment
- ADPretrain: Advancing Industrial Anomaly Detection via Anomaly Representation Pretraining
- EveryDayVLA: A Vision-Language-Action Model for Affordable Robotic Manipulation
- Turning Adversaries into Allies: Reversing Typographic Attacks for Multimodal E-Commerce Product Retrieval
- Multi-modal Loop Closure Detection with Foundation Models in Severely Unstructured Environments
- MUSE: Multi-Scale Dense Self-Distillation for Nucleus Detection and Classification
- Another BRIXEL in the Wall: Towards Cheaper Dense Features
- Towards Cellular-Scale Interpretability in Pathology Foundation Models for Biomarker Assessment
- GSE: Evaluating Sticker Visual Semantic Similarity via a General Sticker Encoder
- Grounding Foundational Vision Models with 3D Human Poses for Robust Action Recognition
- Sublinear iterations can suffice even for DDPMs
- Personalized Image Editing in Text-to-Image Diffusion Models via Collaborative Direct Preference Optimization
- Tracking and Understanding Object Transformations
- UniSplat: Unified Spatio-Temporal Fusion via 3D Latent Scaffolds for Dynamic Driving Scene Reconstruction
- Landslide Hazard Mapping with Geospatial Foundation Models: Geographical Generalizability, Data Scarcity, and Band Adaptability
- MacroNav: Multi-Task Context Representation Learning Enables Efficient Navigation in Unknown Environments
- Covariance Descriptors Meet General Vision Encoders: Riemannian Deep Learning for Medical Image Classification
- MedDChest: A Content-Aware Multimodal Foundational Vision Model for Thoracic Imaging
- ForeRobo: Unlocking Infinite Simulation Data for 3D Goal-driven Robotic Manipulation
- Subsampled Randomized Fourier GaLore for Adapting Foundation Models in Depth-Driven Liver Landmark Segmentation
- Accelerating Physical Property Reasoning for Augmented Visual Cognition
- An Augmentation Overlap Theory of Contrastive Learning
- Image-Intrinsic Priors for Integrated Circuit Defect Detection and Novel Class Discovery via Self-Supervised Learning
- WorldPlanner: Monte Carlo Tree Search and MPC with Action-Conditioned Visual World Models
- GeoCrossBench: Cross-Band Generalization for Remote Sensing
- PLUTO-4: Frontier Pathology Foundation Models
- SE(3)-PoseFlow: Estimating 6D Pose Distributions for Uncertainty-Aware Robotic Manipulation
- Differentiable Hierarchical Visual Tokenization
- In-Context Adaptation of VLMs for Few-Shot Cell Detection in Optical Microscopy
- Learning Interactive World Model for Object-Centric Reinforcement Learning
- SciTextures: Collecting and Connecting Visual Patterns, Models, and Code Across Science and Art
- OmniBrainBench: A Comprehensive Multimodal Benchmark for Brain Imaging Analysis Across Multi-stage Clinical Tasks
- Challenging DINOv3 Foundation Model under Low Inter-Class Variability: A Case Study on Fetal Brain Ultrasound
- VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning
- HumanCrafter: Synergizing Generalizable Human Reconstruction and Semantic 3D Segmentation
- Bayesian model selection and misspecification testing in imaging inverse problems only from noisy and partial measurements
- Image Hashing via Cross-View Code Alignment in the Age of Foundation Models
- Fusion of Multi-scale Heterogeneous Pathology Foundation Models for Whole Slide Image Analysis
- DANCER: Dance ANimation via Condition Enhancement and Rendering with diffusion model
- A Step Toward World Models: A Survey on Robotic Manipulation
- Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
- MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-Experts
- Scaling Image Geo-Localization to Continent Level
- Emu3.5: Native Multimodal Models are World Learners
- A filtering scheme for confocal laser endomicroscopy (CLE)-video sequences for self-supervised learning
- LoCoT2V-Bench: A Benchmark for Long-Form and Complex Text-to-Video Generation
- FullPart: Generating each 3D Part at Full Resolution
- Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail
- Active Learning with Task-Driven Representations for Messy Pools
- Robotic Assistant: Completing Collaborative Tasks with Dexterous Vision-Language-Action Models
- Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization
- FaCT: Faithful Concept Traces for Explaining Neural Network Decisions
- Beyond Facial Consistency: Personalized Person Image Generation with Holistic Identity Preservation
- MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
- SeFi-Image: A Text-to-Image Foundation Model with Semantic-First Diffusion
- How Out-of-Equilibrium Phase Transitions can Seed Pattern Formation in Trained Diffusion Models
- Tissue concepts: Supervised foundation models in computational pathology
- Challenges in data-driven geospatial modeling for environmental research and practice
- A Theory of Contrastive Learning with Natural Images
- CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling
- Anchoring and Steering Diffusion: Enhancing the Faithfulness of Text-to-Image Generation at Inference Time
- Lightweight Image Classification of Raptor Species for Edge Devices: Rare-Species Dataset Expansion via Video Frame Extraction, Knowledge Distillation, and TensorRT Deployment
- Rad-JEPA 3D: Radiology Joint-Embedding Predictive Model for 3D Computed Tomography
- Paris: A Decentralized Trained Open-Weight Diffusion Model
- Amortized Moment Matching for Visual Generation
- GAS-MIL: Group-Aggregative Selection Multi-Instance Learning for Ensemble of Foundation Models in Digital Pathology Image Analysis
- StructureGS: Structure-aware Gaussian Splatting for Articulated Object Reconstruction
- Robust RPC Bundle Adjustment for Multi-Date Satellite Imagery with Season-Invariant Correspondences
- SCALPEL: Semantic Cross-modal Alignment via LLM-Powered Encoder Learning for Medical Vision-Language Representation
- PixARMesh: Autoregressive Mesh-Native Single-View Scene Reconstruction
- Torus embeddings
- Scene-Centric Unsupervised Video Panoptic Segmentation
- SpectraDINO: Modality-Conditioned Adaptation of RGB Vision Foundation Models Across Infrared Bands
- Is Dimensionality a Barrier for Retrieval Models?
- Platonic Representations in the Human Brain: Unsupervised Recovery of Universal Geometry
- Predictive coding video models capture dorsal parietal representations and human judgments for surfaces defined by motion
- Seeking the Unfamiliar but Memorable: Conceptual Creativity as Meta-Learning
- MIND: Monge Inception Distance for Generative Models Evaluation
- Beyond Prompts: Unconditional 3D Inversion for Out-of-Distribution Shapes
- Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
- DINOSim: Zero-Shot Object Detection and Semantic Segmentation on Microscopy Images
- RA-Det: Towards Universal Detection of AI-Generated Images via Robustness Asymmetry
- SPROUT: A Scalable Diffusion Foundation Model for Agricultural Vision
- RL makes MLLMs see better than SFT
- XRefine: Attention-Guided Keypoint Match Refinement
- Attentive multilayer fusion for vision transformers
- Scanner-Induced Domain Shifts Undermine the Robustness of Pathology Foundation Models
- A multimodal whole-slide foundation model for pathology
- Test-Time Adaptive Object Detection with Foundation Model
- Transformers in Medicine: Improving Vision-Language Alignment for Medical Image Captioning
- EA3D: Online Open-World 3D Object Extraction from Streaming Videos
- DINO-YOLO: Self-Supervised Pre-training for Data-Efficient Object Detection in Civil Engineering Applications
- Decoupled MeanFlow: Turning Flow Models into Flow Maps for Accelerated Sampling
- Does Object Binding Naturally Emerge in Large Pretrained Vision Transformers?
- Eigenfunction Extraction for Ordered Representation Learning
- When are radiology reports useful for training medical image classifiers?
- OpenLVLM-MIA: A Controlled Benchmark Revealing the Limits of Membership Inference Attacks on Large Vision-Language Models
- Perception Learning: A Formal Separation of Sensory Representation Learning from Decision Learning
- A Unified Geometric Space Bridging AI Models and the Human Brain
- DynaRend: Learning 3D Dynamics via Masked Future Rendering for Robotic Manipulation
- PRIVET: Privacy Metric Based on Extreme Value Theory
- Kernelized Sparse Fine-Tuning with Bi-level Parameter Competition for Vision Models
- Self-supervised Synthetic Pretraining for Inference of Stellar Mass Embedded in Dense Gas
- Differential Privacy: Gradient Leakage Attacks in Federated Learning Environments
- Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation
- Endowing GPT-4 with a Humanoid Body: Building the Bridge Between Off-the-Shelf VLMs and the Physical World
- Using UAV images and deep learning to enhance the mapping of deadwood in boreal forests
- SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning
- MoS-VLA: A Vision-Language-Action Model with One-Shot Skill Adaptation
- CountFormer: A Transformer Framework for Learning Visual Repetition and Structure in Class-Agnostic Object Counting
- RoboOmni: Proactive Robot Manipulation in Omni-modal Context
- Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations
- A Survey on Efficient Vision-Language-Action Models
- UrbanVLA: A Vision-Language-Action Model for Urban Micromobility
- FreeFuse: Multi-Subject LoRA Fusion via Adaptive Token-Level Routing at Test Time
- Symmetria: A Synthetic Dataset for Learning in Point Clouds
- Model-Behavior Alignment under Flexible Evaluation: When the Best-Fitting Model Isn't the Right One
- ReconViaGen: Towards Accurate Multi-view 3D Object Reconstruction via Generation
- DecoDINO: 3D Human-Scene Contact Prediction with Semantic Classification
- SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene Consistency
- Cross-view Localization and Synthesis -- Datasets, Challenges and Opportunities
- IGGT: Instance-Grounded Geometry Transformer for Semantic 3D Reconstruction
- RoboSVG: A Unified Framework for Interactive SVG Generation with Multi-modal Guidance
- PSScreen V2: Partially Supervised Multiple Retinal Disease Screening
- EndoSfM3D: Learning to 3D Reconstruct Any Endoscopic Surgery Scene using Self-supervised Foundation Model
- Simplifying Knowledge Transfer in Pretrained Models
- CogStereo: Neural Stereo Matching with Implicit Spatial Cognition Embedding
- Sprint: Sparse-Dense Residual Fusion for Efficient Diffusion Transformers
- WorldGrow: Generating Infinite 3D World
- ArtiLatent: Realistic Articulated 3D Object Generation via Structured Latents
- Bridging the gap to real-world language-grounded visual concept learning
- TerraGen: A Unified Multi-Task Layout Generation Framework for Remote Sensing Data Augmentation
- Randomized-MLP Regularization Improves Domain Adaptation and Interpretability in DINOv2
- In Silico Mapping of Visual Categorical Selectivity Across the Whole Brain
- NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation
- Generalizable Hierarchical Skill Learning via Object-Centric Representation
- SafetyPairs: Isolating Safety Critical Image Features with Counterfactual Image Generation
- Controllable-LPMoE: Adapting to Challenging Object Segmentation via Dynamic Local Priors from Mixture-of-Experts
- Modest-Align: Data-Efficient Alignment for Vision-Language Models
- Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
- Blockwise Flow Matching: Improving Flow Matching Models For Efficient High-Quality Generation
- BioDet: Boosting Industrial Object Detection with Image Preprocessing Strategies
- OnlineSplatter: Pose-Free Online 3D Reconstruction for Free-Moving Objects
- HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models
- Why Prototypes Collapse: Diagnosing and Preventing Partial Collapse in Prototypical Self-Supervised Learning
- AlphaFlow: Understanding and Improving MeanFlow Models
- Dino-Diffusion Modular Designs Bridge the Cross-Domain Gap in Autonomous Parking
- Exploring Conditions for Diffusion models in Robotic Control
- Seed3D 1.0: From Images to High-Fidelity Simulation-Ready 3D Assets
- Data-Centric Lessons To Improve Speech-Language Pretraining
- The Intricate Dance of Prompt Complexity, Quality, Diversity, and Consistency in T2I Models
- Using Non-Expert Data to Robustify Imitation Learning via Offline Reinforcement Learning
- Towards Single-Source Domain Generalized Object Detection via Causal Visual Prompts
- Unified Reinforcement and Imitation Learning for Vision-Language Models
- Advances in 4D Representation: Geometry, Motion, and Interaction
- Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks
- X-Ego: Acquiring Team-Level Tactical Situational Awareness via Cross-Egocentric Contrastive Video Representation Learning
- Exploring Scale Shift in Crowd Localization under the Context of Domain Generalization
- GRASPLAT: Enabling dexterous grasping through novel view synthesis
- SEAL: Semantic-Aware Hierarchical Learning for Generalized Category Discovery
- Learning Task-Agnostic Representations through Multi-Teacher Distillation
- CytoNet: A Foundation Model for the Human Cerebral Cortex at Cellular Resolution
- RayPose: Ray Bundling Diffusion for Template Views in Unseen 6D Object Pose Estimation
- Beyond Single Images: Retrieval Self-Augmented Unsupervised Camouflaged Object Detection
- DSI-Bench: A Benchmark for Dynamic Spatial Intelligence
- MoGA: Mixture-of-Groups Attention for End-to-End Long Video Generation
- VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
- RaindropGS: A Benchmark for 3D Gaussian Splatting under Raindrop Conditions
- Towards 3D Objectness Learning in an Open World
- Adapting a global plant identification model to detect invasive alien plant species in high-resolution road side images
- DETree: DEtecting Human-AI Collaborative Texts via Tree-Structured Hierarchical Representation Learning
- M2H: Multi-Task Learning with Efficient Window-Based Cross-Task Attention for Monocular Spatial Perception
- Nearest-Class Mean and Logits Agreement for Wildlife Open-Set Recognition
- Implicit State Estimation via Video Replanning
- Learning After Model Deployment
- AION-1: Omnimodal Foundation Model for Astronomical Sciences
- One Dinomaly2 Detect Them All: A Unified Framework for Full-Spectrum Unsupervised Anomaly Detection
- Exploring Structural Degradation in Dense Representations for Self-supervised Learning
- Xiaoice: Training-Free Video Understanding via Self-Supervised Spatio-Temporal Clustering of Semantic Features
- From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
- Optimizing DINOv2 with Registers for Face Anti-Spoofing
- Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey
- Mapping Hidden Heritage: Self-supervised Pre-training on High-Resolution LiDAR DEM Derivatives for Archaeological Stone Wall Detection
- GaussGym: An open-source real-to-sim framework for learning locomotion from pixels
- CaMiT: A Time-Aware Car Model Dataset for Classification and Generation
- How Universal Are SAM2 Features?
- DINO-CVA: A Multimodal Goal-Conditioned Vision-to-Action Model for Autonomous Catheter Navigation
- Do Satellite Tasks Need Special Pretraining?
- Symmetric Entropy-Constrained Video Coding for Machines
- Universal and Transferable Attacks on Pathology Foundation Models
- DriveGen3D: Boosting Feed-Forward Driving Scene Generation with Efficient Video Diffusion
- Memory-SAM: Human-Prompt-Free Tongue Segmentation via Retrieval-to-Prompt
- VO-DP: Semantic-Geometric Adaptive Diffusion Policy for Vision-Only Robotic Manipulation
- StretchySnake: Flexible SSM Training Unlocks Action Recognition Across Spatio-Temporal Scales
- Comprehensive language-image pre-training for 3D medical image understanding
- Learning an Image Editing Model without Image Editing Pairs
- WithAnyone: Towards Controllable and ID Consistent Image Generation
- Adapting Self-Supervised Representations as a Latent Space for Efficient Generation
- Zero-Shot Wildlife Sorting Using Vision Transformers: Evaluating Clustering and Continuous Similarity Ordering
- Towards Generalist Intelligence in Dentistry: Vision Foundation Models for Oral and Maxillofacial Radiology
- Semantic representations emerge in biologically inspired ensembles of cross-supervising neural networks
- Spatial Preference Rewarding for MLLMs Spatial Understanding
- Vision-Centric Activation and Coordination for Multimodal Large Language Models
- MatchAttention: Matching the Relative Positions for High-Resolution Cross-View Matching
- NANO3D: A Training-Free Approach for Efficient 3D Editing Without Masks
- Salient Concept-Aware Generative Data Augmentation
- UrbanVerse: Scaling Urban Simulation by Watching City-Tour Videos
- Towards Adversarial Robustness and Uncertainty Quantification in DINOv2-based Few-Shot Anomaly Detection
- MUSE: Model-based Uncertainty-aware Similarity Estimation for zero-shot 2D Object Detection and Segmentation
- Exploratory Causal Inference in SAEnce
- Synchronization of Multiple Videos
- Scaling Vision Transformers for Functional MRI with Flat Maps
- DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning
- Visual Interestingness Decoded: How GPT-4o Mirrors Human Interests
- Provenance of AI-Generated Images: A Vector Similarity and Blockchain-based Approach
- PhysMaster: Mastering Physical Representation for Video Generation via Reinforcement Learning
- InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
- The Mechanistic Emergence of Symbol Grounding in Language Models
- Scope: Selective Cross-modal Orchestration of Visual Perception Experts
- DiffEM: Learning from Corrupted Data with Diffusion Models via Expectation Maximization
- AnyUp: Universal Feature Upsampling
- LayerSync: Self-aligning Intermediate Layers
- Unlocking Zero-Shot Plant Segmentation with Pl@ntNet Intelligence
- Unifying Vision-Language Latents for Zero-label Image Caption Enhancement
- Diffusion Transformers with Representation Autoencoders
- Readout Representation: Redefining Neural Codes by Input Recovery
- Actron3D: Learning Actionable Neural Functions from Videos for Transferable Robotic Manipulation
- Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
- Point Prompting: Counterfactual Tracking with Video Diffusion Models
- Scaling Language-Centric Omnimodal Representation Learning
- ARMADA: Autonomous Online Failure Detection and Human Shared Control Empower Scalable Real-world Deployment and Adaptation
- Beyond 'Templates': Category-Agnostic Object Pose, Size, and Shape Estimation from a Single View
- ACE-G: Improving Generalization of Scene Coordinate Regression Through Query Pre-Training
- SCOOP'D: Learning Mixed-Liquid-Solid Scooping via Sim2Real Generative Policy
- How many samples to label for an application given a foundation model? Chest X-ray classification study
- HiMaCon: Discovering Hierarchical Manipulation Concepts from Unlabeled Multi-Modal Data
- InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models
- Neural Weight Compression for Language Models
- G2L:From Giga-Scale to Cancer-Specific Large-Scale Pathology Foundation Models via Knowledge Distillation
- MSCloudCAM: Multi-Scale Context Adaptation with Convolutional Cross-Attention for Multispectral Cloud Segmentation
- Enhancing Zero-Shot Anomaly Detection: CLIP-SAM Collaboration with Cascaded Prompts
- Prompt-Guided Spatial Understanding with RGB-D Transformers for Fine-Grained Object Relation Reasoning
- Visual Odometry with Transformers
- Structured Spectral Graph Representation Learning for Multi-label Abnormality Analysis from 3D CT Scans
- Equipping Vision Foundation Model with Mixture of Experts for Out-of-Distribution Detection
- UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
- Understanding Self-supervised Contrastive Learning through Supervised Objectives
- BitMar: Low-Bit Multimodal Fusion with Episodic Memory for Edge Devices
- Unified Open-World Segmentation with Multi-Modal Prompts
- On the Problem of Consistent Anomalies in Zero-Shot Industrial Anomaly Detection
- FRIEREN: Federated Learning with Vision-Language Regularization for Segmentation
- SoundReactor: Frame-level Online Video-to-Audio Generation
- Vision4PPG: Emergent PPG Analysis Capability of Vision Foundation Models for Vital Signs like Blood Pressure
- Bridging Perspectives: Foundation Model Guided BEV Maps for 3D Object Detection and Tracking
- VG-Mapping: Variation-Aware 3D Gaussians for Online Semi-static Scene Mapping
- From Generic to Specialized: A Subspecialty Diagnostic System Powered by Self-Supervised Learning for Cervical Histopathology
- J-RAS: Mutual Adaptation for Medical Image Segmentation via Contrastive Retrieval-Augmented Joint Optimization
- AFFORD2ACT: Affordance-Guided Automatic Keypoint Selection for Generalizable and Lightweight Robotic Manipulation
- DreamX-World 1.0: A General-Purpose Interactive World Model
- Leveraging Prior Knowledge of Diffusion Model for Person Search
- SSeg: Active Sparse Point-Label Augmentation for Semantic Segmentation
- Holistic Order Prediction in Natural Scenes
- Myopic Bayesian Decision Theory for Batch Active Learning with Partial Batch Label Sampling
- Hybrid-grained Feature Aggregation with Coarse-to-fine Language Guidance for Self-supervised Monocular Depth Estimation
- Towards Safer and Understandable Driver Intention Prediction
- Visual Anomaly Detection for Reliable Robotic Implantation of Flexible Microelectrode Array
- D-TPT: Dimensional Entropy Maximization for Calibrating Test-Time Prompt Tuning in Vision-Language Models
- Vision Language Models: A Survey of 26K Papers
- ReSplat: Learning Recurrent Gaussian Splats
- To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
- Evaluating Small Vision-Language Models on Distance-Dependent Traffic Perception
- Latent Harmony: Synergistic Unified UHD Image Restoration via Latent Space Regularization and Controllable Refinement
- XYZCylinder: Towards Compatible Feed-Forward 3D Gaussian Splatting for Driving Scenes via Unified Cylinder Lifting Method
- Self-Supervised Learning Strategies for a Platform to Test the Toxicity of New Chemicals and Materials
- AlignGS: Aligning Geometry and Semantics for Robust Indoor Reconstruction from Sparse Views
- SyncHuman: Synchronizing 2D and 3D Generative Models for Single-view Human Reconstruction
- InstructUDrag: Joint Text Instructions and Object Dragging for Interactive Image Editing
- NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
- Re-Identifying Kākā with AI-Automated Video Key Frame Extraction
- Unlocking 3D Affordance Segmentation with 2D Semantic Knowledge
- Don't Run with Scissors: Pruning Breaks VLA Models but They Can Be Recovered
- Beyond Textual CoT: Interleaved Text-Image Chains with Deep Confidence Reasoning for Image Editing
- Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers
- Resolution scaling governs DINOv3 transfer performance in chest radiograph classification
- TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking
- DADO: A Depth-Attention framework for Object Discovery
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- Revisiting Mixout: An Overlooked Path to Robust Finetuning
- VA-Adapter: Adapting Ultrasound Foundation Model to Echocardiography Probe Guidance
- UniFField: A Generalizable Unified Neural Feature Field for Visual, Semantic, and Spatial Uncertainties in Any Scene
- Heptapod: Language Modeling on Visual Signals
- SDQM: Synthetic Data Quality Metric for Object Detection Dataset Evaluation
- Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
- VUGEN: Visual Understanding priors for GENeration
- SIGMA-GEN: Structure and Identity Guided Multi-subject Assembly for Image Generation
- EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark
- Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
- Shaken or Stirred? An Analysis of MetaFormer's Token Mixing for Medical Imaging
- Kaputt: A Large-Scale Dataset for Visual Defect Detection
- Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow
- Redefining Generalization in Visual Domains: A Two-Axis Framework for Fake Image Detection with FusionDetect
- Efficient Conditional Generation on Scale-based Visual Autoregressive Models
- nnSAM2: nnUNet-Enhanced One-Prompt SAM2 for Few-shot Multi-Modality Segmentation and Composition Analysis of Lumbar Paraspinal Muscles
- Human3R: Everyone Everywhere All at Once
- Information-Theoretic Policy Pre-Training with Empowerment
- Midway Network: Learning Representations for Recognition and Motion from Latent Dynamics
- Boomerang Distillation Enables Zero-Shot Model Size Interpolation
- Factuality Matters: When Image Generation and Editing Meet Structured Visuals
- StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation
- SegMASt3R: Geometry Grounded Segment Matching
- SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
- HyperVLA: Efficient Inference in Vision-Language-Action Models via Hypernetworks
- Activation Quantization of Vision Encoders Needs Prefixing Registers
- Foundation Visual Encoders Are Secretly Few-Shot Anomaly Detectors
- REAR: Rethinking Visual Autoregressive Models via Generator-Tokenizer Consistency Regularization
- Domain Generalization Under Posterior Drift
- Pulp Motion: Framing-aware multimodal camera and human motion generation
- Mitigating Diffusion Model Hallucinations with Dynamic Guidance
- VER: Vision Expert Transformer for Robot Learning via Foundation Distillation and Dynamic Routing
- Glocal Information Bottleneck for Time Series Imputation
- ActiveMark: on watermarking of visual foundation models via massive activations
- Scaling Sequence-to-Sequence Generative Neural Rendering
- Visual Lifelog Retrieval through Captioning-Enhanced Interpretation
- Flexible and Efficient Spatio-Temporal Transformer for Sequential Visual Place Recognition
- Seeing the Bigger Picture: 3D Latent Mapping for Mobile Manipulation Policy Learning
- The Overlooked Value of Test-time Reference Sets in Visual Place Recognition
- Mapping Rio de Janeiro's favelas: general-purpose vs. satellite-specific neural networks
- EmbodiSwap for Zero-Shot Robot Imitation Learning
- A Hybrid Co-Finetuning Approach for Visual Bug Detection in Video Games
- FrameOracle: Learning What to See and How Much to See in Videos
- Efficient Surgical Robotic Instrument Pose Reconstruction in Real World Conditions Using Unified Feature Detection
- Geometry Meets Vision: Revisiting Pretrained Semantics in Distilled Fields
- When and Where do Events Switch in Multi-Event Video Generation?
- Towards Scalable and Consistent 3D Editing
- Representing Beauty: Towards a Participatory but Objective Latent Aesthetics
- One Patch to Caption Them All: A Unified Zero-Shot Captioning Framework
- Brain-Language Model Alignment: Insights into the Platonic Hypothesis and Intermediate-Layer Advantage
- KeySG: Hierarchical Keyframe-Based 3D Scene Graphs
- TD-JEPA: Latent-predictive Representations for Zero-Shot Reinforcement Learning
- Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
- ProtoMask: Segmentation-Guided Prototype Learning
- Robust Context-Aware Object Recognition
- Normal-Abnormal Guided Generalist Anomaly Detection
- VLOD-TTA: Test-Time Adaptation of Vision-Language Object Detectors
- Selective Underfitting in Diffusion Models
- BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration
- MorphGen: Controllable and Morphologically Plausible Generative Cell-Imaging
- DiSC-AMC: Token- and Parameter-Efficient Discretized Statistics In-Context Automatic Modulation Classification
- Stitch: Training-Free Position Control in Multimodal Diffusion Transformers
- Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
- Are neural scaling laws leading quantum chemistry astray?
- MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation
- Seeing Space and Motion: Enhancing Latent Actions with Spatial and Dynamic Awareness for VLA
- Towards Continual Expansion of Data Coverage: Automatic Text-guided Edge-case Synthesis
- EchoGen: Generating Visual Echoes in Any Scene via Feed-Forward Subject-Driven Auto-Regressive Model
- EasyOcc: 3D Pseudo-Label Supervision for Fully Self-Supervised Semantic Occupancy Prediction Models
- DGM4+: Dataset Extension for Global Scene Inconsistency
- SGS: Segmentation-Guided Scoring for Global Scene Inconsistencies
- CO3: Contrasting Concepts Compose Better
- PatchEAD: Unifying Industrial Visual Prompting Frameworks for Patch-Exclusive Anomaly Detection
- RAE: A Neural Network Dimensionality Reduction Method for Nearest Neighbors Preservation in Vector Search
- Adapting SAM with Dynamic Similarity Graphs for Few-Shot Parameter-Efficient Small Dense Object Detection: A Case Study of Chickpea Pods in Field Conditions
- Best of Sim and Real: Decoupled Visuomotor Manipulation via Learning Control in Simulation and Perception in Real
- SAGE: Spatial-visual Adaptive Graph Exploration for Visual Place Recognition
- Towards Reliable and Holistic Visual In-Context Learning Prompt Selection
- TTT3R: 3D Reconstruction as Test-Time Training
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- Mitigating Hallucination in Multimodal LLMs with Layer Contrastive Decoding
- BRIDGE -- Building Reinforcement-Learning Depth-to-Image Data Generation Engine for Monocular Depth Estimation
- Annotation-Free One-Shot Imitation Learning for Multi-Step Manipulation Tasks
- ThermalGen: Style-Disentangled Flow-Based Generative Models for RGB-to-Thermal Image Translation
- UP2You: Fast Reconstruction of Yourself from Unconstrained Photo Collections
- Vision Function Layer in Multimodal LLMs
- SkyLink: Unifying Street-Satellite Geo-Localization via UAV-Mediated 3D Scene Alignment
- IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video?
- DINOReg: Strong Point Cloud Registration with Vision Foundation Model
- Similarity-Aware Selective State-Space Modeling for Semantic Correspondence
- Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers
- Towards Foundation Models for Cryo-ET Subtomogram Analysis
- Combining Discrepancy-Confusion Uncertainty and Calibration Diversity for Active Fine-Grained Image Classification
- Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification
- Fast Feature Field (F3): A Predictive Representation of Events
- World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training
- Fidelity-Aware Data Composition for Robust Robot Generalization
- NeoWorld: Neural Simulation of Explorable Virtual Worlds via Progressive 3D Unfolding
- Mitigating Visual Hallucinations via Semantic Curriculum Preference Optimization in MLLMs
- Personalized Vision via Visual In-Context Learning
- Scalable GANs with Transformers
- UniLat3D: Geometry-Appearance Unified Latents for Single-Stage 3D Generation
- SCOPE: Semantic Conditioning for Sim2Real Category-Level Object Pose Estimation in Robotics
- PAD3R: Pose-Aware Dynamic 3D Reconstruction from Casual Videos
- Does Weak-to-strong Generalization Happen under Spurious Correlations?
- RPG360: Robust 360 Depth Estimation with Perspective Foundation Models and Graph Optimization
- AnyDepth: Depth Estimation Made Easy
- Revisit the Imbalance Optimization in Multi-task Learning: An Experimental Analysis
- Towards Fine-Grained Text-to-3D Quality Assessment: A Benchmark and A Two-Stage Rank-Learning Metric
- HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
- Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models
- Color-Pair Guided Robust Zero-Shot 6D Pose Estimation and Tracking of Cluttered Objects on Edge Devices
- ZeroScene: A Zero-Shot Framework for 3D Scene Generation from a Single Image and Controllable Texture Editing
- VMDiff: Visual Mixing Diffusion for Limitless Cross-Object Synthesis
- Deep Taxonomic Networks for Unsupervised Hierarchical Prototype Discovery
- Generalizable Coarse-to-Fine Robot Manipulation via Language-Aligned 3D Keypoints
- Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensional
- Multi-Modal Manipulation via Multi-Modal Policy Consensus
- WeatherCycle: Unpaired Multi-Weather Restoration via Color Space Decoupled Cycle Learning
- GLUE: Global-Local Unified Encoding for Imitation Learning via Key-Patch Tracking
- Unsupervised Online 3D Instance Segmentation with Synthetic Sequences and Dynamic Loss
- Transferring Vision-Language-Action Models to Industry Applications: Architectures, Performance, and Challenges
- Streamline pathology foundation model by cross-magnification distillation
- Understanding Language Prior of LVLMs by Contrasting Chain-of-Embedding
- Confidence-Calibrating Regularization for Robust Brain MRI Segmentation Under Domain Shift
- MindCraft: How Concept Trees Take Shape In Deep Models
- PartCo: Part-Level Correspondence Priors Enhance Category Discovery
- WoW: Towards a World omniscient World model Through Embodied Interaction
- Orochi: Versatile Biomedical Image Processor
- Category Discovery: An Open-World Perspective
- Learning the Neighborhood: Contrast-Free Multimodal Self-Supervised Molecular Graph Pretraining
- RAU: Reference-based Anatomical Understanding with Vision Language Models
- ReLAM: Learning Anticipation Model for Rewarding Visual Robotic Manipulation
- UniMapGen: A Generative Framework for Large-Scale Map Construction from Multi-modal Data
- Polysemous Language Gaussian Splatting via Matching-based Mask Lifting
- Action-aware Dynamic Pruning for Efficient Vision-Language-Action Manipulation
- MultiCrafter: High-Fidelity Multi-Subject Generation via Disentangled Attention and Identity-Aware Preference Alignment
- DynaNav: Dynamic Feature and Layer Selection for Efficient Visual Navigation
- SingRef6D: Monocular Novel Object Pose Estimation with a Single RGB Reference
- PANICL: Mitigating Over-Reliance on Single Prompt in Visual In-Context Learning
- Drag4D: Align Your Motion with Text-Driven 3D Scene Generation
- VLBiMan: Vision-Language Anchored One-Shot Demonstration Enables Generalizable Bimanual Robotic Manipulation
- On the Status of Foundation Models for SAR Imagery
- PartSAM: A Scalable Promptable Part Segmentation Model Trained on Native 3D Data
- SlotFM: A Motion Foundation Model with Slot Attention for Diverse Downstream Tasks
- Temporal vs. Spatial: Comparing DINOv3 and V-JEPA2 Feature Representations for Video Action Analysis
- What Happens Next? Anticipating Future Motion by Generating Point Trajectories
- No Alignment Needed for Generation: Learning Linearly Separable Representations in Diffusion Models
- Unsupervised Defect Detection for Surgical Instruments
- SD3.5-Flash: Distribution-Guided Distillation of Generative Flows
- Quantized Visual Geometry Grounded Transformer
- Does FLUX Already Know How to Perform Physically Plausible Image Composition?
- Every Subtlety Counts: Fine-grained Person Independence Micro-Action Recognition via Distributionally Robust Optimization
- Decipher-MR: A Vision-Language Foundation Model for 3D MRI Representations
- RetoVLA: Reusing Register Tokens for Spatial Reasoning in Vision-Language-Action Models
- SiNGER: A Clearer Voice Distills Vision Transformers Further
- Dense Semantic Matching with VGGT Prior
- The Unanticipated Asymmetry Between Perceptual Optimization and Assessment
- FSMODNet: A Closer Look at Few-Shot Detection in Multispectral Data
- FreeInsert: Personalized Object Insertion with Geometric and Style Control
- Large Pre-Trained Models for Bimanual Manipulation in 3D
- Anomaly Detection by Clustering DINO Embeddings using a Dirichlet Process Mixture
- GS-RoadPatching: Inpainting Gaussians via 3D Searching and Placing for Driving Scenes
- Latent Iterative Refinement Flow: A Geometric-Constrained Approach for Few-Shot Generation
- CAMILA: Context-Aware Masking for Image Editing with Language Alignment
- Enhancing Transformer-Based Vision Models: Addressing Feature Map Anomalies Through Novel Optimization Strategies
- Hyperspectral Adapter for Semantic Segmentation with Vision Foundation Models
- GraspFactory: A Large Object-Centric Grasping Dataset
- AnySafe: Adapting Latent Safety Filters at Runtime via Safety Constraint Parameterization in the Latent Space
- VCP-DCN: Beyond Visual Concealed Property via Depth Collaborative Network for Camouflaged Object Detection
- World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models
- RefineSVG: Visual Feedback-Driven Reinforcement Learning for Image-to-SVG Generation
- Foundation Models for Face Presentation Attack Detection: A Unified Linear-Probing Benchmark
- MeshFM: 2D Features Are All You Need for 3D Shape Understanding
- SiamJEPA: On the Role of Siamese Student Encoders in JEPA
- Beyond Classification: Pathology Foundation Models as Detection Encoders for Mitotic Figures
- MUL-T: Decoding Spatial Cellular Architecture in Multiplexed Tissue Images
- DexDirect: Direct Kinesthetic Arm Guidance for Efficient Dexterous Demonstration Collection
- RFMSR: Residual Flow Matching for Image Super-Resolution
- What Makes Deep Learning Work for Traditional Chinese Medicine Tongue Diagnosis? A Comprehensive Ablation Study
- Split and Drive: Dual-Axis Disentanglement for Real-Time Gaussian Head Avatars
- ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow
- Towards Real-Time PixOOD: Efficient Anomaly Segmentation for Autonomous Vehicles
- VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding
- Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer
- Failure Detection for Surgical Robot Imitation Policies via Flow-Matching World Modeling
- AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes
- Driving on Registers
- A scalar per patch from pre-trained ViTs enables fast moving navigation in the real world
- 3D-Fixer: Coarse-to-Fine In-place Completion for 3D Scenes from a Single Image
- Using reflectance spectra and Pl@ ntNet to identify herbarium specimens: a case study with Lithocarpus
- Temporal Straightening for Latent Planning
- Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis
- Dark3R: Learning Structure from Motion in the Dark
- DAGE: Dual-Stream Architecture for Efficient and Fine-Grained Geometry Estimation
- OSF: On Pre-training and Scaling of Sleep Foundation Models
- RAD: Retrieval-Augmented Monocular Metric Depth Estimation for Underrepresented Classes
- SAIL-Recon: Large SfM by Augmenting Scene Regression with Localization
- EndoUFM: Utilizing Foundation Models for Monocular depth estimation of endoscopic images
- Seeing Like a Designer Without One: A Study on Unsupervised Slide Quality Assessment via Designer Cue Augmentation
- The Platonic Universe: Do Foundation Models See the Same Sky?
- ConViS-Bench: Estimating Video Similarity Through Semantic Concepts
- ManipForce: Force-Guided Policy Learning with Frequency-Aware Representation for Contact-Rich Manipulation
- Advancing Metallic Surface Defect Detection via Anomaly-Guided Pretraining on a Large Industrial Dataset
- How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
- Zero-shot Monocular Metric Depth for Endoscopic Images
- VGGT-DP: Generalizable Robot Control via Vision Foundation Models
- The Photographer Eye: Teaching Multimodal Large Language Models to Understand Image Aesthetics like Photographers
- Attack for Defense: Adversarial Agents for Point Prompt Optimization Empowering Segment Anything Model
- FUNCanon: Learning Pose-Aware Action Primitives via Functional Object Canonicalization for Generalizable Robotic Manipulation
- 3D Flow Diffusion Policy: Visuomotor Policy Learning via Generating Flow in 3D Space
- Latent Action Pretraining Through World Modeling
- Enhancing Semantic Segmentation with Continual Self-Supervised Pre-training
- SSNet: Flexible and robust channel extrapolation for fluid antenna systems enabled by an self-supervised learning framework
- EigenSafe: A Spectral Framework for Learning-Based Probabilistic Safety Assessment
- Development and validation of an AI foundation model for endoscopic diagnosis of esophagogastric junction adenocarcinoma: a cohort and deep learning study
- VideoArtGS: Building Digital Twins of Articulated Objects from Monocular Video
- Overview of PlantCLEF 2025: Multi-Species Plant Identification in Vegetation Quadrat Images
- Visual Instruction Pretraining for Domain-Specific Foundation Models
- Training-Free Label Space Alignment for Universal Domain Adaptation
- Revisiting Vision Language Foundations for No-Reference Image Quality Assessment
- Evict3R: Training-Free Token Eviction for Memory-Bounded Streaming Visual Geometry Transformers
- ProDyG: Progressive Dynamic Scene Reconstruction via Gaussian Splatting from Monocular Videos
- Prepare Before You Act: Learning From Humans to Rearrange Initial States
- MRN: Harnessing 2D Vision Foundation Models for Diagnosing Parkinson's Disease with Limited 3D MR Data
- SPFSplatV2: Efficient Self-Supervised Pose-Free 3D Gaussian Splatting from Sparse Views
- Catching the Details: Self-Distilled RoI Predictors for Fine-Grained MLLM Perception
- Parameter-efficient fine-tuning (PEFT) of Vision Foundation Models for Atypical Mitotic Figure Classification
- HOGraspFlow: Exploring Vision-based Generative Grasp Synthesis with Hand-Object Priors and Taxonomy Awareness
- Random Direct Preference Optimization for Radiography Report Generation
- Informative Text-Image Alignment for Visual Affordance Learning with Foundation Models
- DiffEye: Diffusion-Based Continuous Eye-Tracking Data Generation Conditioned on Natural Images
- Mixture of Noise for Pre-Trained Model-Based Class-Incremental Learning
- A Closer Look at Model Collapse: From a Generalization-to-Memorization Perspective
- KV-Efficient VLA: A Method to Speed up Vision Language Models with RNN-Gated Chunked KV Cache
- \boldsymbolλ-Orthogonality Regularization for Compatible Representation Learning
- OS-DiffVSR: Towards One-step Latent Diffusion Model for High-detailed Real-world Video Super-Resolution
- IDU: Incremental Dynamic Update of Existing 3D Virtual Environments with New Imagery Data
- SegDINO3D: 3D Instance Segmentation Empowered by Both Image-Level and Object-Level 2D Features
- DistillMatch: Leveraging Knowledge Distillation from Vision Foundation Model for Multimodal Image Matching
- Overview of PlantCLEF 2024: multi-species plant identification in vegetation plot images
- UNIV: Unified Foundation Model for Infrared and Visible Modalities
- Minimal Semantic Sufficiency Meets Unsupervised Domain Generalization
- MS-GS: Multi-Appearance Sparse-View 3D Gaussian Splatting in the Wild
- Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
- RangeSAM: On the Potential of Visual Foundation Models for Range-View represented LiDAR segmentation
- Emulating Human-like Adaptive Vision for Efficient and Flexible Machine Visual Perception
- Depth AnyEvent: A Cross-Modal Distillation Paradigm for Event-Based Monocular Depth Estimation
- RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
- Trade-offs in Cross-Domain Generalization of Foundation Model Fine-Tuned for Biometric Applications
- Which Direction to Choose? An Analysis on the Representation Power of Self-Supervised ViTs in Downstream Tasks
- FMGS-Avatar: Mesh-Guided 2D Gaussian Splatting with Foundation Model Priors for 3D Monocular Avatar Reconstruction
- Autoguided Online Data Curation for Diffusion Model Training
- Designing Latent Safety Filters using Pre-Trained Vision Models
- DACoN: DINO for Anime Paint Bucket Colorization with Any Number of Reference Images
- Toward Embodiment Equivariant Vision-Language-Action Policy
- OpenViGA: Video Generation for Automotive Driving Scenes by Streamlining and Fine-Tuning Open Source Models with Public Data
- AToken: A Unified Tokenizer for Vision
- Hashing-Baseline: Rethinking Hashing in the Age of Pretrained Models
- AssoCiAm: A Benchmark for Evaluating Association Thinking while Circumventing Ambiguity
- An Exploratory Study on Abstract Images and Visual Representations Learned from Them
- Data Leakage in Visual Datasets
- Distractor-Aware Memory-Based Visual Object Tracking
- Neural Proteomics Fields for Super-resolved Spatial Proteomics Prediction
- MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment
- MetricNet: Recovering Metric Scale in Generative Navigation Policies
- Pre-Manipulation Alignment Prediction with Parallel Deep State-Space and Transformer Models
- Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation
- CSMoE: An Efficient Remote Sensing Foundation Model with Soft Mixture-of-Experts
- Consistent View Alignment Improves Foundation Models for 3D Medical Image Segmentation
- Reinforcement Learning for Robotic Insertion of Flexible Cables in Industrial Settings
- FishBEV: Distortion-Resilient Bird's Eye View Segmentation with Surround-View Fisheye Cameras
- Learning Discrete Abstractions for Visual Rearrangement Tasks Using Vision-Guided Graph Coloring
- Gaussian Alignment for Relative Camera Pose Estimation via Single-View Reconstruction
- Diving into Mitigating Hallucinations from a Vision Perspective for Large Vision-Language Models
- First Place Solution to the MLCAS 2025 GWFSS Challenge: The Devil is in the Detail and Minority
- From Embeddings to Equations: Genetic-Programming Surrogates for Interpretable Transformer Classification
- Brought a Gun to a Knife Fight: Modern VFM Baselines Outgun Specialized Detectors on In-the-Wild AI Image Detection
- SHREC 2025: Protein surface shape retrieval including electrostatic potential
- MEGAN: Mixture of Experts for Robust Uncertainty Estimation in Endoscopy Videos
- NavMoE: Hybrid Model- and Learning-based Traversability Estimation for Local Navigation via Mixture of Experts
- SPGen: Spherical Projection as Consistent and Flexible Representation for Single Image 3D Shape Generation
- Case-Based Decision-Theoretic Decoding with Quality Memories
- MFAF: An EVA02-Based Multi-scale Frequency Attention Fusion Method for Cross-View Geo-Localization
- Beyond Averages: Open-Vocabulary 3D Scene Understanding with Gaussian Splatting and Bag of Embeddings
- Data Scaling Laws for Radiology Foundation Models
- From reactive to cognitive: brain-inspired spatial intelligence for embodied agents
- The Better You Learn, The Smarter You Prune: Towards Efficient Vision-language-action Models via Differentiable Token Pruning
- Pre-trained Visual Representations Generalize Where it Matters in Model-Based Reinforcement Learning
- MMMS: Multi-Modal Multi-Surface Interactive Segmentation
- Multi Anatomy X-Ray Foundation Model
- DS@GT AnimalCLEF: Triplet Learning over ViT Manifolds with Nearest Neighbor Classification for Animal Re-identification
- Character-Centric Understanding of Animated Movies
- LoRA-fine-tuned Large Vision Models for Automated Assessment of Post-SBRT Lung Injury
- Embodied Navigation Foundation Model
- FS-SAM2: Adapting Segment Anything Model 2 for Few-Shot Semantic Segmentation via Low-Rank Adaptation
- A Computer Vision Pipeline for Individual-Level Behavior Analysis: Benchmarking on the Edinburgh Pig Dataset
- RAM++: Robust Representation Learning via Adaptive Mask for All-in-One Image Restoration
- DRAG: Data Reconstruction Attack using Guided Diffusion
- RAPTOR: A Foundation Policy for Quadrotor Control
- Synthetic vs. Real Training Data for Visual Navigation
- Domain-Adaptive Pretraining Improves Primate Behavior Recognition
- Tenma: Robust Cross-Embodiment Robot Manipulation with Diffusion Transformer
- DinoAtten3D: Slice-Level Attention Aggregation of DinoV2 for 3D Brain MRI Anomaly Classification
- SSG-Dit: A Spatial Signal Guided Framework for Controllable Video Generation
- Beyond Instance Consistency: Investigating View Diversity in Self-supervised Learning
- Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
- Leveraging Geometric Priors for Unaligned Scene Change Detection
- UnLoc: Leveraging Depth Uncertainties for Floorplan Localization
- Self-supervised Learning Of Visual Pose Estimation Without Pose Labels By Classifying LED States
- GLAM: Geometry-Guided Local Alignment for Multi-View VLP in Mammography
- LayerLock: Non-collapsing Representation Learning with Progressive Freezing
- BenchECG and xECG: a benchmark and baseline for ECG foundation models
- Loc2: Interpretable Cross-View Localization via Depth-Lifted Local Feature Matching
- Unified Multimodal Model as Auto-Encoder
- Graph Alignment via Dual-Pass Spectral Encoding and Latent Space Communication
- Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders
- Unsupervised Integrated-Circuit Defect Segmentation via Image-Intrinsic Normality
- VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
- Learning Object-Centric Representations in SAR Images with Multi-Level Feature Fusion
- MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos
- Diffusion-Based Action Recognition Generalizes to Untrained Domains
- Vision-Language Semantic Aggregation Leveraging Foundation Model for Generalizable Medical Image Segmentation
- ViewSparsifier: Killing Redundancy in Multi-View Plant Phenotyping
- Chirality in Action: Time-Aware Video Representation Learning by Latent Straightening
- Rethinking the Backbone in Class Imbalanced Federated Source Free Domain Adaptation: The Utility of Vision Foundation Models
- Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities
- RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation
- Visual Representation Alignment for Multimodal Large Language Models
- Three Pillars improving Vision Foundation Model Distillation for Lidar
- Domain Knowledge is Power: Leveraging Physiological Priors for Self Supervised Representation Learning in Electrocardiography
- RINO: Renormalization Group Invariance with No Labels
- PanoLAM: Large Avatar Model for Gaussian Full-Head Synthesis from One-shot Unposed Image
- P3-SAM: Native 3D Part Segmentation
- Online Clustering of Seafloor Imagery for Interpretation during Long-Term AUV Operations
- MM-DINOv2: Adapting Foundation Models for Multi-Modal Medical Image Analysis
- Does DINOv3 Set a New Medical Vision Standard? Benchmarking 2D and 3D Classification, Segmentation, and Registration
- Block-Sparse Global Attention for Efficient Multi-View Geometry Transformers
- O3Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation
- Investigating Location-Regularised Self-Supervised Feature Learning for Seafloor Visual Imagery
- Curia: A Multi-Modal Foundation Model for Radiology
- SVGauge: Towards Human-Aligned Evaluation for SVG Generation
- UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning
- S-LAM3D: Segmentation-Guided Monocular 3D Object Detection via Feature Space Fusion
- Closer to Reality: Practical Semi-Supervised Federated Learning for Foundation Model Adaptation
- PictOBI-20k: Unveiling Large Multimodal Models in Visual Decipherment for Pictographic Oracle Bone Characters
- Patch-Level Kernel Alignment for Dense Self-Supervised Learning
- Reconstruction and Reenactment Separated Method for Realistic Gaussian Head
- SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning
- Prior Distribution and Model Confidence
- Missing Fine Details in Images: Last Seen in High Frequencies
- SGS-3D: High-Fidelity 3D Instance Segmentation via Reliable Semantic Mask Splitting and Growing
- LUIVITON: Learned Universal Interoperable VIrtual Try-ON
- Dynamic Group Detection using VLM-augmented Temporal Groupness Graph
- Towards Open World Detection: A Survey
- Symbolic Graphics Programming with Large Language Models
- Few-step Flow for 3D Generation via Marginal-Data Transport Distillation
- FPC-VLA: A Vision-Language-Action Framework with a Supervisor for Failure Prediction and Correction
- Weakly-Supervised Learning of Dense Functional Correspondences
- Reactive In-Air Clothing Manipulation with Confidence-Aware Dense Correspondence and Visuotactile Affordance
- Efficient Odd-One-Out Anomaly Detection
- Global-to-Local or Local-to-Global? Enhancing Image Retrieval with Efficient Local Search and Effective Global Re-ranking
- Differentiable Entropy Regularization: A Complexity-Aware Approach for Neural Optimization
- HAMSt3R: Human-Aware Multi-view Stereo 3D Reconstruction
- Scalable and Loosely-Coupled Multimodal Deep Learning for Breast Cancer Subtyping
- DUViN: Diffusion-Based Underwater Visual Navigation via Knowledge-Transferred Depth Features
- Planning with Reasoning using Vision Language World Model
- Toward a robust lesion detection model in breast DCE-MRI: adapting foundation models to high-risk women
- FastVGGT: Training-Free Acceleration of Visual Geometry Transformer
- Motion-Refined DINOSAUR for Unsupervised Multi-Object Discovery
- Manipulation as in Simulation: Enabling Accurate Geometry Perception in Robots
- Vision encoders should be image size agnostic and task driven
- SALAD -- Semantics-Aware Logical Anomaly Detection
- MOSAIC: Multi-Subject Personalized Generation via Correspondence-Aware Alignment and Disentanglement
- An Investigation of Visual Foundation Models Robustness
- DIET-CP: Lightweight and Data Efficient Self Supervised Continued Pretraining
- 2D Gaussian Splatting with Semantic Alignment for Image Inpainting
- T-MASK: Temporal Masking for Probing Foundation Models across Camera Views in Driver Monitoring
- First RAG, Second SEG: A Training-Free Paradigm for Camouflaged Object Detection
- MedDINOv3: How to adapt vision foundation models for medical image segmentation?
- Examination of PCA Utilisation for Multilabel Classifier of Multispectral Images
- FEDEXCHANGE: Bridging the Domain Gap in Federated Object Detection for Free
- Improving Large Vision and Language Models by Learning from a Panel of Peers
- Take That for Me: Multimodal Exophora Resolution with Interactive Questioning for Ambiguous Out-of-View Instructions
- SoccerHigh: A Benchmark Dataset for Automatic Soccer Video Summarization
- M3Ret: Unleashing Zero-shot Multimodal Medical Image Retrieval via Self-Supervision
- ReCap: Event-Aware Image Captioning with Article Retrieval and Semantic Gaussian Normalization
- EndoGMDE: Generalizable Monocular Depth Estimation with Mixture of Low-Rank Experts for Diverse Endoscopic Scenes
- FocusDPO: Dynamic Preference Optimization for Multi-Subject Personalized Image Generation via Adaptive Focus
- Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation
- Advances and Trends in the 3D Reconstruction of the Shape and Motion of Animals
- MoTo: A Zero-shot Plug-in Interaction-aware Navigation for General Mobile Manipulation
- Correlates of Image Memorability in Vision Encoders: Activations, Attention Entropy, Patch Uniformity and Autoencoder Losses
- Prompt the Unseen: Evaluating Visual-Language Alignment Beyond Supervision
- Decomposing and Revising What Language Models Generate
- SegDINO: An Efficient Design for Medical and Natural Image Segmentation with DINO-V3
- HERO-VQL: Hierarchical, Egocentric and Robust Visual Query Localization
- Adaptive Point-Prompt Tuning: Fine-Tuning Heterogeneous Foundation Models for 3D Point Cloud Analysis
- Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment
- Category-level Text-to-Image Retrieval Improved: Bridging the Domain Gap with Diffusion Models and Vision Encoders
- TMUAD: Enhancing Logical Capabilities in Unified Anomaly Detection Models with a Text Memory Bank
- OptMark: Robust Multi-bit Diffusion Watermarking via Inference Time Optimization
- Federated Fine-tuning of SAM-Med3D for MRI-based Dementia Classification
- MM-SeR: Multimodal Self-Refinement for Lightweight Image Captioning
- Generative AI for Industrial Contour Detection: A Language-Guided Vision System
- Representation Learning with Adaptive Superpixel Coding
- Maybe you don't need a U-Net: convolutional feature upsampling for materials micrograph segmentation
- PHD: Personalized 3D Human Body Fitting with Point Diffusion
- Generalizable Object Re-Identification via Visual In-Context Prompting
- Native Logical and Hierarchical Representations with Subspace Embeddings
- Prompt-to-Product: Generative Assembly via Bimanual Manipulation
- Webly-Supervised Image Manipulation Localization via Category-Aware Auto-Annotation
- Dino U-Net: Exploiting High-Fidelity Dense Features from Foundation Models for Medical Image Segmentation
- Adapting Foundation Model for Dental Caries Detection with Dual-View Co-Training
- SceneGen: Single-Image 3D Scene Generation in One Feedforward Pass
- Waver: Wave Your Way to Lifelike Video Generation
- StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
- Disentangling Latent Embeddings with Sparse Linear Concept Subspaces (SLiCS)
- Plug-in Feedback Self-adaptive Attention in CLIP for Training-free Open-Vocabulary Segmentation
- AudioStory: Generating Long-Form Narrative Audio with Large Language Models
- Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
- Self-supervised structured object representation learning
- FastAvatar: Towards Unified Fast High-Fidelity 3D Avatar Reconstruction with Large Gaussian Reconstruction Transformers
- IELDG: Suppressing Domain-Specific Noise with Inverse Evolution Layers for Domain Generalized Semantic Segmentation
- Interact-Custom: Customized Human Object Interaction Image Generation
- DNP-Guided Contrastive Reconstruction with a Reverse Distillation Transformer for Medical Anomaly Detection
- MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
- PRISM: A Framework Harnessing Unsupervised Visual Representations and Textual Prompts for Explainable MACE Survival Prediction from Cardiac Cine MRI
- LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry Grounding
- Toward Robust Medical Fairness: Debiased Dual-Modal Alignment via Text-Guided Attribute-Disentangled Prompt Learning for Vision-Language Models
- eSkinHealth: A Multimodal Dataset for Neglected Tropical Skin Diseases
- SAT-SKYLINES: 3D Building Generation from Satellite Imagery and Coarse Geometric Priors
- CLARIFY: A Specialist-Generalist Framework for Accurate and Lightweight Dermatological Visual Question Answering
- Follow My Hold: Hand-Object Interaction Reconstruction through Geometric Guidance
- Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
- SpotEdit: Evaluating Visually-Guided Image Editing Methods
- Annotation-Free Open-Vocabulary Segmentation for Remote-Sensing Images
- Scaling Group Inference for Diverse and High-Quality Generation
- RCDINO: Enhancing Radar-Camera 3D Object Detection with DINOv2 Semantic Features
- Survey of Vision-Language-Action Models for Embodied Manipulation
- Multi-Object Sketch Animation with Grouping and Motion Trajectory Priors
- Image-Conditioned 3D Gaussian Splat Quantization
- Deep learning in plant phenotyping: the first ten years
- TransLight: Image-Guided Customized Lighting Control with Generative Decoupling
- Controllable Latent Space Augmentation for Digital Pathology
- NoteIt: A System Converting Instructional Videos to Interactable Notes Through Multimodal Video Understanding
- SignBind-LLM: Multi-Stage Modality Fusion for Sign Language Translation
- DINOv3 with Test-Time Training for Medical Image Registration
- Tinker: Diffusion's Gift to 3D--Multi-View Consistent Editing From Sparse Inputs without Per-Scene Optimization
- Seeing Further on the Shoulders of Giants: Knowledge Inheritance for Vision Foundation Models
- GALA: Guided Attention with Language Alignment for Open Vocabulary Gaussian Splatting
- Local Scale Equivariance with Latent Deep Equilibrium Canonicalizer
- GeoSAM2: Unleashing the Power of SAM2 for 3D Part Segmentation
- Backdooring Self-Supervised Contrastive Learning by Noisy Alignment
- Self-Supervised Sparse Sensor Fusion for Long Range Perception
- RCGNet: RGB-based Category-Level 6D Object Pose Estimation with Geometric Guidance
- Breaking the SFT Plateau: Multimodal Structured Reinforcement Learning for Chart-to-Code Generation
- MimicFunc: Imitating Tool Manipulation from a Single Human Video via Functional Correspondence
- Calibrating Biased Distribution in VFM-derived Latent Space via Cross-Domain Geometric Consistency
- Revisiting MLLM Token Technology through the Lens of Classical Visual Coding
- Unleashing Semantic and Geometric Priors for 3D Scene Completion
- Sim-to-Real Dynamic Object Manipulation on Conveyor Systems via Optimization Path Shaping
- Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- Learn Faster and Remember More: Balancing Exploration and Exploitation for Continual Test-time Adaptation
- Manipulate-to-Navigate: Reinforcement Learning with Visual Affordances and Manipulability Priors
- MATPAC++: Enhanced Masked Latent Prediction for Self-Supervised Audio Representation Learning
- IGFuse: Interactive 3D Gaussian Scene Reconstruction via Multi-Scans Fusion
- DermINO: Hybrid Pretraining for a Versatile Dermatology Foundation Model
- Splat Feature Solver
- S5: Scalable Semi-Supervised Semantic Segmentation in Remote Sensing
- Towards interpretable prediction of recurrence risk in breast cancer using pathology foundation models
- UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding
- Human Centric General Physical Intelligence for Agile Manufacturing Automation
- CoFi: A Fast Coarse-to-Fine Few-Shot Pipeline for Glomerular Basement Membrane Segmentation
- MM-R1: Unleashing the Power of Unified Multimodal Large Language Models for Personalized Image Generation
- Remove360: Benchmarking Residuals After Object Removal in 3D Gaussian Splatting
- UniDCF: A Foundation Model for Comprehensive Dentocraniofacial Hard Tissue Reconstruction
- Probing the Representational Power of Sparse Autoencoders in Vision Models
- Generalized Decoupled Learning for Enhancing Open-Vocabulary Dense Perception
- CoreEditor: Correspondence-constrained Diffusion for Consistent 3D Editing
- StyleMM: Stylized 3D Morphable Face Model via Text-Driven Aligned Image Translation
- VFM-Guided Semi-Supervised Detection Transformer under Source-Free Constraints for Remote Sensing Object Detection
- FusionFM: Fusing Eye-specific Foundational Models for Optimized Ophthalmic Diagnosis
- Visual Perception Engine: Fast and Flexible Multi-Head Inference for Robotic Vision Tasks
- MAESTRO: Masked AutoEncoders for Multimodal, Multitemporal, and Multispectral Earth Observation Data
- Privacy-Aware Detection of Fake Identity Documents: Methodology, Benchmark, and Improved Algorithms (FakeIDet2)
- ViewBridge:Revisiting Cross-View Localization from Image Matching
- Processing and acquisition traces in visual encoders: What does CLIP know about your camera?
- Increasing the Utility of Synthetic Images through Chamfer Guidance
- Multi-Label Plant Species Prediction with Metadata-Enhanced Multi-Head Vision Transformers
- Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
- Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
- PQ-DAF: Pose-driven Quality-controlled Data Augmentation for Data-scarce Driver Distraction Detection
- Towards Efficient Prompt-based Continual Learning in Distributed Medical AI
- From Pixel to Mask: A Survey of Out-of-Distribution Segmentation
- SynBrain: Enhancing Visual-to-fMRI Synthesis via Probabilistic Representation Learning
- DINOMotion: advanced robust tissue motion tracking with DINOv2 in 2D-Cine MRI-guided radiotherapy
- Deep Learning for Crack Detection: A Review of Learning Paradigms, Generalizability, and Datasets
- Story2Board: A Training-Free Approach for Expressive Storyboard Generation
- A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and Generation
- Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion Models
- SHREC'25 Track on Multiple Relief Patterns: Report and Analysis
- GSFixer: Improving 3D Gaussian Splatting with Reference-Guided Video Diffusion Priors
- MInDI-3D: Iterative Deep Learning in 3D for Sparse-view Cone Beam Computed Tomography
- Gen-AFFECT: Generation of Avatar Fine-grained Facial Expressions with Consistent identiTy
- SkySplat: Generalizable 3D Gaussian Splatting from Multi-Temporal Sparse Satellite Images
- RASR: Retrieval-Augmented Super Resolution for Practical Reference-based Image Restoration
- Masquerade: Learning from In-the-wild Human Videos using Data-Editing
- Collaborative Face Experts Fusion in Video Generation: Boosting Identity Consistency Across Large Face Poses
- MADPromptS: Unlocking Zero-Shot Morphing Attack Detection with Multiple Prompt Aggregation
- Exploring Palette based Color Guidance in Diffusion Models
- Scaling Learned Image Compression Models up to 1 Billion
- Cut2Next: Generating Next Shot via In-Context Tuning
- RedDino: A foundation model for red blood cell analysis
- ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model
- TRIDE: A Text-assisted Radar-Image weather-aware fusion network for Depth Estimation
- From Field to Drone: Domain Drift Tolerant Automated Multi-Species and Damage Plant Semantic Segmentation for Herbicide Trials
- Splat4D: Diffusion-Enhanced 4D Gaussian Splatting for Temporally and Spatially Consistent Content Creation
- Anatomy-Aware Low-Dose CT Denoising via Pretrained Vision Models and Semantic-Guided Contrastive Learning
- Comparison Reveals Commonality: Customized Image Generation through Contrastive Inversion
- LET-US: Long Event-Text Understanding of Scenes
- Pref-GUIDE: Continual Policy Learning from Real-Time Human Feedback via Preference-Based Learning
- BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
- A Semi-Supervised Learning Method for the Identification of Bad Exposures in Large Imaging Surveys
- CLIPin: A Non-contrastive Plug-in to CLIP for Multimodal Semantic Alignment
- Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation
- SAM Encoder Breach by Adversarial Simplicial Complex Triggers Downstream Model Failures
- Bounding Distributional Shifts in World Modeling through Novelty Detection
- SwiftVideo: A Unified Framework for Few-Step Video Generation through Trajectory-Distribution Alignment
- ThematicPlane: Bridging Tacit User Intent and Latent Spaces for Image Generation
- Toward Context-Aware Exoskeleton Assistance: Integrating Computer Vision Payload Estimation with a User-Centric Optimization Space
- AnomalyMoE: Towards a Language-free Generalist Model for Unified Visual Anomaly Detection
- Robust Image Stitching with Optimal Plane
- Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
- Hi3DEval: Advancing 3D Generation Evaluation with Hierarchical Validity
- Symmetry Understanding of 3D Shapes via Chirality Disentanglement
- SMOL-MapSeg: Show Me One Label as prompt
- EndoMatcher: Generalizable Endoscopic Image Matcher via Multi-Domain Pre-training for Robot-Assisted Surgery
- AdaFusion: Prompt-Guided Inference with Adaptive Fusion of Pathology Foundation Models
- A Study of the Framework and Real-World Applications of Language Embedding for 3D Scene Understanding
- Propagating Sparse Depth via Depth Foundation Model for Out-of-Distribution Depth Completion
- AdvDINO: Domain-Adversarial Self-Supervised Representation Learning for Spatial Proteomics
- Voost: A Unified and Scalable Diffusion Transformer for Bidirectional Virtual Try-On and Try-Off
- Open Scene Graphs for Open-World Object-Goal Navigation
- Perch 2.0: The Bittern Lesson for Bioacoustics
- BridgeDepth: Bridging Monocular and Stereo Reasoning with Latent Alignment
- Visual Bias and Interpretability in Deep Learning for Dermatological Image Analysis
- One Model for All: Unified Try-On and Try-Off in Any Pose via LLM-Inspired Bidirectional Tweedie Diffusion
- Composed Object Retrieval: Object-level Retrieval via Composed Expressions
- Towards Resilient Safety-driven Unlearning for Diffusion Models against Downstream Fine-tuning
- InfoQ: Mixed-Precision Quantization via Global Information Flow
- Conditional Latent Diffusion Models for Zero-Shot Instance Segmentation
- Radar-Based NLoS Pedestrian Localization for Darting-Out Scenarios Near Parked Vehicles with Camera-Assisted Point Cloud Interpretation
- Constraint-Preserving Data Generation for Visuomotor Policy Learning
- Statistical Confidence Rescoring for Robust 3D Scene Graph Generation from Multi-View Images
- Uncertainty-aware Accurate Elevation Modeling for Off-road Navigation via Neural Processes
- Prediction-Oriented Subsampling from Data Streams
- Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition
- OmniShape: Zero-Shot Multi-Hypothesis Shape and Pose Estimation in the Real World
- MetaScope: Optics-Driven Neural Network for Ultra-Micro Metalens Endoscopy
- SAM2-UNeXT: An Improved High-Resolution Baseline for Adapting Foundation Models to Downstream Segmentation Tasks
- MAUP: Training-free Multi-center Adaptive Uncertainty-aware Prompting for Cross-domain Few-shot Medical Image Segmentation
- Prototype-Enhanced Confidence Modeling for Cross-Modal Medical Image-Report Retrieval
- MedCAL-Bench: A Comprehensive Benchmark on Cold-Start Active Learning with Foundation Models for Medical Image Analysis
- SlotMatch: Distilling Object-Centric Representations for Unsupervised Video Segmentation
- Zero-shot Shape Classification of Nanoparticles in SEM Images using Vision Foundation Models
- H3R: Hybrid Multi-view Correspondence for Generalizable 3D Reconstruction
- ADSeeker: A Knowledge-Infused Framework for Anomaly Detection and Reasoning
- Uni3R: Unified 3D Reconstruction and Semantic Understanding via Generalizable Gaussian Splatting from Unposed Multi-View Images
- Multi-Granularity Feature Calibration via VFM for Domain Generalized Semantic Segmentation
- FedPromo: Federated Lightweight Proxy Models at the Edge Bring New Domains to Foundation Models
- D2PPO: Diffusion Policy Policy Optimization with Dispersive Loss
- SGAD: Semantic and Geometric-aware Descriptor for Local Feature Matching
- GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting
- Unified Category-Level Object Detection and Pose Estimation from RGB Images using 3D Prototypes
- RICL: Adding In-Context Adaptability to Pre-Trained Vision-Language-Action Models
- Self-Supervised YOLO: Leveraging Contrastive Learning for Label-Efficient Object Detection
- Mapillary Vistas Validation for Fine-Grained Traffic Signs: A Benchmark Revealing Vision-Language Model Limitations
- Versatile Transition Generation with Image-to-Video Diffusion
- Rein++: Efficient Generalization and Adaptation for Semantic Segmentation with Vision Foundation Models
- Enhancing Zero-Shot Brain Tumor Subtype Classification via Fine-Grained Patch-Text Alignment
- EvoVLMA: Evolutionary Vision-Language Model Adaptation
- Measuring and Predicting Where and When Pathologists Focus their Visual Attention while Grading Whole Slide Images of Cancer
- Set Pivot Learning: Redefining Generalized Segmentation with Vision Foundation Models
- Beyond Vulnerabilities: A Survey of Adversarial Attacks as Both Threats and Defenses in Computer Vision Systems
- StrandDesigner: Towards Practical Strand Generation with Sketch Guidance
- M3LLM: Model Context Protocol-aided Mixture of Vision Experts For Multimodal LLMs in Networks
- Foundation Models for Bioacoustics -- a Comparative Review
- OCSplats: Observation Completeness Quantification and Label Noise Separation in 3DGS
- A Large-Scale Benchmark of Cross-Modal Learning for Histology and Gene Expression in Spatial Transcriptomics
- COLLAGE: Adaptive Fusion-based Retrieval for Augmented Policy Learning
- GECO: Geometrically Consistent Embedding with Lightspeed Inference
- Can Large Pretrained Depth Estimation Models Help With Image Dehazing?
- Is Tracking really more challenging in First Person Egocentric Vision?
- IAMAP: Unlocking Deep Learning in QGIS for non-coders and limited computing resources
- Uncertainty-Aware Likelihood Ratio Estimation for Pixel-Wise Out-of-Distribution Detection
- CoProU-VO: Combining Projected Uncertainty for End-to-End Unsupervised Monocular Visual Odometry
- Training-Free Class Purification for Open-Vocabulary Semantic Segmentation
- Video Color Grading via Look-Up Table Generation
- SDMatte: Grafting Diffusion Models for Interactive Matting
- Representation Shift: Unifying Token Compression with FlashAttention
- LAMIC: Layout-Aware Multi-Image Composition via Scalability of Multimodal Diffusion Transformer
- MVHybrid: Improving Spatial Transcriptomics Prediction with Hybrid State Space-Vision Transformer Backbone in Pathology Vision Foundation Models
- Robust Classification under Noisy Labels: A Geometry-Aware Reliability Framework for Foundation Models
- Uncovering Latent Connections in Indigenous Heritage: Semantic Pipelines for Cultural Preservation in Brazil
- Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis
- SeqAffordSplat: Scene-level Sequential Affordance Reasoning on 3D Gaussian Splatting
- Efficient Masked Attention Transformer for Few-Shot Classification and Segmentation
- DivControl: Knowledge Diversion for Controllable Image Generation
- Gaussian Splatting Feature Fields for Privacy-Preserving Visual Localization
- 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection
- Fusion of Pervasive RF Data with Spatial Images via Vision Transformers for Enhanced Mapping in Smart Cities
- π3: Permutation-Equivariant Visual Geometry Learning
- Mitigating Resolution-Drift in Federated Learning: Case of Keypoint Detection
- iLRM: An Iterative Large 3D Reconstruction Model
- PixNerd: Pixel Neural Field Diffusion
- Zero-Shot Retrieval for Scalable Visual Search in a Two-Sided Marketplace
- CHROMA: Consistent Harmonization of Multi-View Appearance via Bilateral Grid Prediction
- A Real-Time System for Egocentric Hand-Object Interaction Detection in Industrial Domains
- Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath Prediction
- DepR: Depth Guided Single-view Scene Reconstruction with Instance-level Diffusion
- MoCHA: Advanced Vision-Language Reasoning with MoE Connector and Hierarchical Group Attention
- LOTS of Fashion! Multi-Conditioning for Image Generation via Sketch-Text Pairing
- Towards Blind Bitstream-corrupted Video Recovery via a Visual Foundation Model-driven Framework
- UAVScenes: A Multi-Modal Dataset for UAVs
- TRIBE: TRImodal Brain Encoder for whole-brain fMRI response prediction
- TESPEC: Temporally-Enhanced Self-Supervised Pretraining for Event Cameras
- Temporally Consistent Unsupervised Segmentation for Mobile Robot Perception
- AI in Agriculture: A Survey of Deep Learning Techniques for Crops, Fisheries and Livestock
- From Seeing to Experiencing: Scaling Navigation Foundation Models with Reinforcement Learning
- VeS: Teaching Pixels to Listen Without Supervision
- Foundation Models and Transformers for Anomaly Detection: A Survey
- Control Copy-Paste: Controllable Diffusion-Based Augmentation Method for Remote Sensing Few-Shot Object Detection
- CLEVER: Stream-based Active Learning for Robust Semantic Perception from Human Instructions
- DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation
- MoDeSuite: Robot Learning Task Suite for Benchmarking Mobile Manipulation with Deformable Objects
- Meta CLIP 2: A Worldwide Scaling Recipe
- ObjectGS: Object-aware Scene Reconstruction and Scene Understanding via Gaussian Splatting
- PixelNav: Towards Model-based Vision-Only Navigation with Topological Graphs
- Ensemble Foreground Management for Unsupervised Object Discovery
- Compositional Video Synthesis by Temporal Object-Centric Learning
- FMimic: Foundation Models are Fine-grained Action Learners from Human Videos
- ZSE-Cap: A Zero-Shot Ensemble for Image Retrieval and Prompt-Guided Captioning
- Enhancing Spatial Reasoning through Visual and Textual Thinking
- DAViD: Data-efficient and Accurate Vision Models from Synthetic Data
- Can Foundation Models Predict Fitness for Duty?
- ModalFormer: Multimodal Transformer for Low-Light Image Enhancement
- MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation
- SAMwave: Wavelet-Driven Feature Enrichment for Effective Adaptation of Segment Anything Model
- AnimeColor: Reference-based Animation Colorization with Diffusion Transformers
- Second Competition on Presentation Attack Detection on ID Card
- A Survey of Token Compression for Efficient Multimodal Large Language Models
- VESPA: Towards un(Human)supervised Open-World Pointcloud Labeling for Autonomous Driving
- Exact Reformulation and Optimization for Direct Metric Optimization in Binary Imbalanced Classification
- Region-based Cluster Discrimination for Visual Representation Learning
- CLASP: General-Purpose Clothes Manipulation with Semantic Keypoints
- HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly
- Latent Denoising Makes Good Tokenizers
- Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers
- Hybrid Deep Learning and Handcrafted Feature Fusion for Mammographic Breast Cancer Classification
- Taking Language Embedded 3D Gaussian Splatting into the Wild
- SCALAR: Scale-wise Controllable Visual Autoregressive Learning
- SurgPIS: Surgical-instrument-level Instances and Part-level Semantics for Weakly-supervised Part-aware Instance Segmentation
- Back to the Features: DINO as a Foundation for Video World Models
- Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception
- Adversarial Distribution Matching for Diffusion Distillation Towards Efficient Image and Video Synthesis
- TTS-VAR: A Test-Time Scaling Framework for Visual Auto-Regressive Generation
- Q-Former Autoencoder: A Modern Framework for Medical Anomaly Detection
- LMM-Det: Make Large Multimodal Models Excel in Object Detection
- ViGText: Deepfake Image Detection with Vision-Language Model Explanations and Graph Neural Networks
- LEAF: Latent Diffusion with Efficient Encoder Distillation for Aligned Features in Medical Image Segmentation
- Identifying Prompted Artist Names from Generated Images
- BleedOrigin: Dynamic Bleeding Source Localization in Endoscopic Submucosal Dissection via Dual-Stage Detection and Tracking
- Explaining How Visual, Textual and Multimodal Encoders Share Concepts
- Object segmentation in the wild with foundation models: application to vision assisted neuro-prostheses for upper limbs
- DSFormer: A Dual-Scale Cross-Learning Transformer for Visual Place Recognition
- Descrip3D: Enhancing Large Language Model-based 3D Scene Understanding with Object-Level Text Descriptions
- Image Generators are Generalist Vision Learners
- Hierarchical Planning with Latent World Models
- VOID: Video Object and Interaction Deletion
- DiSCO-3D : Discovering and segmenting Sub-Concepts from Open-vocabulary queries in NeRF
- OptiCorNet: Optimizing Sequence-Based Context Correlation for Visual Place Recognition
- Masked Depth Modeling for Spatial Perception
- PAT++: a cautionary tale about generative visual augmentation for Object Re-identification
- Generalist Forecasting with Frozen Video Models via Latent Diffusion
- Shared representations in brains and models reveal a two-route cortical organization during scene perception
- Safety Certification in the Latent space using Control Barrier Functions and World Models
- Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
- SkySense V2: A Unified Foundation Model for Multi-modal Remote Sensing
- Towards Facilitated Fairness Assessment of AI-based Skin Lesion Classifiers Through GenAI-based Image Synthesis
- CNS-Bench: Benchmarking Image Classifier Robustness Under Continuous Nuisance Shifts
- ViRN: Variational Inference and Distribution Trilateration for Long-Tailed Continual Representation Learning
- PIG-Nav: Key Insights for Pretrained Image Goal Navigation Models
- Ultra3D: Efficient and High-Fidelity 3D Generation with Part Attention
- SADA: Stability-guided Adaptive Diffusion Acceleration
- Content-based 3D Image Retrieval and a ColBERT-inspired Re-ranking for Tumor Flagging and Staging
- VLA-Touch: Enhancing Vision-Language-Action Models with Dual-Level Tactile Feedback
- Not All Starting Points Are Equal: Pre-trained Priors and Their Outsized Impact on Person Identification
- Attention (as Discrete-Time Markov) Chains
- Towards Human-level Intelligence via Human-like Whole-Body Manipulation
- InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
- Hallucination Score: Towards Mitigating Hallucinations in Generative Image Super-Resolution
- ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning
- FISHER: A Foundation Model for Multi-Modal Industrial Signal Comprehensive Representation
- Augmented Reality in Cultural Heritage: A Dual-Model Pipeline for 3D Artwork Reconstruction
- Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling
- HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation
- AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation
- AutoPartGen: Autogressive 3D Part Generation and Discovery
- Revisiting Reliability in the Reasoning-based Pose Estimation Benchmark
- DINO-VO: A Feature-based Visual Odometry Leveraging a Visual Foundation Model
- FashionPose: Text to Pose to Relight Image Generation for Personalized Fashion Visualization
- MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing
- LanePerf: a Performance Estimation Framework for Lane Detection
- UniLGL: Learning Uniform Place Recognition for FOV-limited/Panoramic LiDAR Global Localization
- AD-GS: Object-Aware B-Spline Gaussian Splatting for Self-Supervised Autonomous Driving
- Unsupervised Part Discovery via Descriptor-Based Masked Image Restoration with Optimized Constraints
- Spatial Frequency Modulation for Semantic Segmentation
- Compositional Discrete Latent Code for High Fidelity, Productive Diffusion Models
- Unsupervised Monocular 3D Keypoint Discovery from Multi-View Diffusion Priors
- Prototypical Progressive Alignment and Reweighting for Generalizable Semantic Segmentation
- MonoMVSNet: Monocular Priors Guided Multi-View Stereo Network
- GeoDistill: Geometry-Guided Self-Distillation for Weakly Supervised Cross-View Localization
- NavComposer: Composing Language Instructions for Navigation Trajectories through Action-Scene-Object Modularization
- Are Vision Foundation Models Ready for Out-of-the-Box Medical Image Registration?
- TRAN-D: 2D Gaussian Splatting-based Sparse-view Transparent Object Depth Reconstruction via Physics Simulation for Scene Update
- EEG Foundation Models: A Critical Review of Current Progress and Future Directions
- KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model
- Towards Depth Foundation Model: Recent Trends in Vision-Based Depth Estimation
- Personalized OVSS: Understanding Personal Concept in Open-Vocabulary Semantic Segmentation
- Streaming 4D Visual Geometry Transformer
- Sparse Fine-Tuning of Transformers for Generative Tasks
- Quantize-then-Rectify: Efficient VQ-VAE Training
- Graph World Model
- GT-Loc: Unifying When and Where in Images Through a Joint Embedding Space
- Navigating the Challenges of AI-Generated Image Detection in the Wild: What Truly Matters?
- Reprogramming Vision Foundation Models for Spatio-Temporal Forecasting
- Self-supervised Learning on Camera Trap Footage Yields a Strong Universal Face Embedder
- MoVieS: Motion-Aware 4D Dynamic View Synthesis in One Second
- Task Priors: Enhancing Model Evaluation by Considering the Entire Space of Downstream Tasks
- Prompt2DEM: High-Resolution DEMs for Urban and Open Environments from Global Prompts Using a Monocular Foundation Model
- Self-supervised pretraining of vision transformers for animal behavioral analysis and neural encoding
- Disentanglement and Assessment of Shortcuts in Ophthalmological Retinal Imaging Exams
- Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding
- Calibrated and Robust Foundation Models for Vision-Language and Medical Image Tasks Under Distribution Shift
- PoseLLM: Enhancing Language-Guided Human Pose Estimation with MLP Alignment
- Scaling Laws for Optimal Data Mixtures
- Learning Diffusion Models with Flexible Representation Guidance
- Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation
- Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
- InstaScene: Towards Complete 3D Instance Decomposition and Reconstruction from Cluttered Scenes
- Subject-Consistent and Pose-Diverse Text-to-Image Generation
- Raptor: Scalable Train-Free Embeddings for 3D Medical Volumes Leveraging Pretrained 2D Foundation Models
- Transfer Learning and Mixup for Fine-Grained Few-Shot Fungi Classification
- PanMatch: Unleashing the Potential of Large Vision Models for Unified Matching Models
- Seg-Wild: Interactive Segmentation based on 3D Gaussian Splatting for Unconstrained Image Collections
- AI Should Sense Better, Not Just Scale Bigger: Adaptive Sensing as a Paradigm Shift
- NexViTAD: Few-shot Unsupervised Cross-Domain Defect Detection via Vision Foundation Models and Multi-Task Learning
- Where are we with calibration under dataset shift in image classification?
- Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling
- ConsNoTrainLoRA: Data-driven Weight Initialization of Low-rank Adapters using Constraints
- ADIEE: Automatic Dataset Creation and Scorer for Instruction-Guided Image Editing Evaluation
- Combining Pre-Trained Models for Enhanced Feature Representation in Reinforcement Learning
- GreenHyperSpectra: A multi-source hyperspectral dataset for global vegetation trait prediction
- Text-promptable Object Counting via Quantity Awareness Enhancement
- EXAONE Path 2.0: Pathology Foundation Model with End-to-End Supervision
- Image Can Bring Your Memory Back: A Novel Multi-Modal Guided Attack against Image Generation Model Unlearning
- Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor
- Adaptive Part Learning for Fine-Grained Generalized Category Discovery: A Plug-and-Play Enhancement
- Does Data Scaling Lead to Visual Compositional Generalization?
- MK-Pose: Category-Level Object Pose Estimation via Multimodal-Based Keypoint Learning
- Perception-Aware Policy Optimization for Multimodal Reasoning
- AnthroTAP: Learning Point Tracking with Real-World Motion
- Feed-Forward SceneDINO for Unsupervised Semantic Scene Completion
- Is Diversity All You Need for Scalable Robotic Manipulation?
- Improving Robustness of Foundation Models in Domain Adaptation with Soup-Adapters
- NeoBabel: A Multilingual Open Tower for Visual Generation
- Tile-Based ViT Inference with Visual-Cluster Priors for Zero-Shot Multi-Species Plant Identification
- Inter- and Intra-image Refinement for Few Shot Segmentation
- DreamArt: Generating Interactable Articulated Objects from a Single Image
- Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation
- Accelerating Diffusion Model Training under Minimal Budgets: A Condensation-Based Perspective
- Tora2: Motion and Appearance Customized Diffusion Transformer for Multi-Entity Video Generation
- OTAS: Open-vocabulary Token Alignment for Outdoor Segmentation
- Structured Task Solving via Modular Embodied Intelligence: A Case Study on Rubik's Cube
- OmniPart: Part-Aware 3D Generation with Semantic Decoupling and Structural Cohesion
- Generative Panoramic Image Stitching
- Beyond One Shot, Beyond One Perspective: Cross-View and Long-Horizon Distillation for Better LiDAR Representations
- Weighted Ensemble Models Are Strong Continual Learners
- ConBatch-BAL: Batch Bayesian Active Learning under Budget Constraints
- Leveraging Self-Supervised Features for Efficient Flooded Region Identification in UAV Aerial Images
- From Imitation to Innovation: The Emergence of AI Unique Artistic Styles and the Challenge of Copyright Protection
- VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs
- Learning Robust Stereo Matching in the Wild with Selective Mixture-of-Experts
- A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation
- X-ray transferable polyrepresentation learning
- A View-consistent Sampling Method for Regularized Training of Neural Radiance Fields
- DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
- ZERO: Industry-ready Vision Foundation Model with Multi-modal Prompts
- Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation
- VICI: VLM-Instructed Cross-view Image-localisation
- Query-Based Adaptive Aggregation for Multi-Dataset Joint Training Toward Universal Visual Place Recognition
- On the rankability of visual embeddings
- SciVid: Cross-Domain Evaluation of Video Models in Scientific Applications
- Foundation versus Domain-specific Models: Performance Comparison, Fusion, and Explainability in Face Recognition
- PhenoBench: A Comprehensive Benchmark for Cell Phenotyping
- Zero-shot Inexact CAD Model Alignment from a Single Image
- Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
- LACONIC: A 3D Layout Adapter for Controllable Image Creation
- FastDINOv2: Frequency Based Curriculum Learning Improves Robustness and Training Speed
- LiteReality: Graphics-Ready 3D Scene Reconstruction from RGB-D Scans
- Linear Attention with Global Context: A Multipole Attention Mechanism for Vision and Physics
- APT: Adaptive Personalized Training for Diffusion Models with Limited Data
- Embedding-Based Federated Data Sharing via Differentially Private Conditional VAEs
- MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp Details
- PLOT: Pseudo-Labeling via Video Object Tracking for Scalable Monocular 3D Object Detection
- AvatarMakeup: Realistic Makeup Transfer for 3D Animatable Head Avatars
- HyperGaussians: High-Dimensional Gaussian Splatting for High-Fidelity Animatable Face Avatars
- IC-Custom: Diverse Image Customization via In-Context Learning
- Enhanced Generative Model Evaluation with Clipped Density and Coverage
- A Gift from the Integration of Discriminative and Diffusion-based Generative Learning: Boundary Refinement Remote Sensing Semantic Segmentation
- MARVIS: Modality Adaptive Reasoning over VISualizations
- BronchoGAN: Anatomically consistent and domain-agnostic image-to-image translation for video bronchoscopy
- Vision transformer [wikipedia]
- Geospatial foundation model [wikipedia]
- Reverse image search [wikipedia]
- Vision–language–action model [wikipedia]
Discussions
- dino v2 so your model takes an image as input, and outputs a vector the objective is that if you crop and rotate the image at random, it outputs the same vector that's it this turns out to extract ton [bsky, 122 points, 9 comments]
- AFAIK it's the same dataset, they just use the larger pretrained model as the teacher model. Screenshot is from the DinoV2 paper section 5: arxiv.org/abs/2304.07193 [bsky, 1 points, 1 comments]
- distills DINOv2 and SAM features arxiv.org/abs/2304.07193, arxiv.org/abs/2304.02643 [bsky, 1 points, 1 comments]
- Open Source model fromMeta, DINOv2: ViT model, trained with 1 billion parameters and subsequently distilled into a range of smaller models, outperforms OpenCLIP in some benchmarks. arxiv.org/abs/2304. [bsky, 1 points, 1 comments]
- DINO v2 (https://arxiv.org/abs/2304.07193) has learned about object parts: heads, wings are recognized across distinct categories such as birds and planes! 100% unsupervised! (context: see my DINO sel [bsky, 0 points, 1 comments]
Related