V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
2026/03/15 by Lorenzo Mur-Labadia, Matthew J. Muckley, Matthew Muckley +7 · 2 voices · 2 citations
Computer Science · Engineering · #Action recognition #Anticipation (artificial intelligence) #Deep learning #Encoder #Feature (linguistics) #Human Pose and Action Recognition #Key (lock) #Multimodal Machine Learning Applications #Representation (politics) #Robot Manipulation and Learning #Training (meteorology) #cs.CV
paper · pdf · open access · doi:10.48550/arxiv.2603.14482
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2026/03/15 · arxiv published 2026/03/15 · openalex created_date 2026/03/18 · arxiv updated 2026/06/11 · openalex updated_date 2026/07/28
Abstract
We present V-JEPA 2.1, a family of self-supervised models that learn dense, high-quality visual representations for both images and videos while retaining strong global scene understanding. The approach combines four key components. First, a dense predictive loss uses a masking-based objective in which both visible and masked tokens contribute to the training signal, encouraging explicit spatial and temporal grounding. Second, deep self-supervision applies the self-supervised objective hierarchically across multiple intermediate encoder layers to improve representation quality. Third, multi-modal tokenizers enable unified training across images and videos. Finally, the model benefits from effective scaling in both model capacity and training data. Together, these design choices produce representations that are spatially structured, semantically coherent, and temporally consistent. Empirically, V-JEPA 2.1 achieves state-of-the-art performance on several challenging benchmarks, including 7.71 mAP on Ego4D for short-term object-interaction anticipation and 40.8 Recall@5 on EPIC-KITCHENS for high-level action anticipation, as well as a 20-point improvement in real-robot grasping success rate over V-JEPA-2 AC. The model also demonstrates strong performance in robotic navigation (5.687 ATE on TartanDrive), depth estimation (0.307 RMSE on NYUv2 with a linear probe), and global recognition (77.7 on Something-Something-V2). These results show that V-JEPA 2.1 significantly advances the state of the art in dense visual understanding and world modeling.
Citations
- Chimère Ω — blueprint for a physico-cognitively inspired local-first LLM runtime
- DINOv3
- Back to the Features: DINO as a Foundation for Video World Models
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs
- PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
- Perception Encoder: The best visual embeddings are not at the output of the network
- Scaling Language-Free Visual Representation Learning
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Qwen2.5-VL Technical Report
- Intuitive physics understanding emerges from self-supervised pretraining on natural videos
- An Empirical Study of Autoregressive Pre-training from Videos
- Scaling 4D Representations
- DINO-Foresight: Looking into the Future with DINO
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- Navigation World Models
- TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
- Connecting Joint-Embedding Predictive Architecture with Contrastive Self-supervised Learning
- TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- Near, far: Patch-ordering enhances vision foundation models' scene understanding
- No Train, all Gain: Self-Supervised Gradients Improve Deep Frozen Representations
- AFF-ttention! Affordances and Attention models for Short-Term Object Interaction Anticipation
- Can't make an Omelette without Breaking some Eggs: Plausible Action Anticipation using Large Video-Language Models
- InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
- TempCompass: Do Video LLMs Really Understand Videos?
- VideoPrism: A Foundational Visual Encoder for Video Understanding
- Revisiting Feature Prediction for Learning Visual Representations from Video
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data
- AM-RADIO: Agglomerative Vision Foundation Model -- Reduce All Domains Into One
- MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
- Vision Transformers Need Registers
- Time Does Tell: Self-Supervised Time-Tuning of Dense Image Representations
- Stochastic positional embeddings improve masked image modeling
- Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
- Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles
- Perception Test: A Diagnostic Benchmark for Multimodal Video Models
- DINOv2: Learning Robust Visual Features without Supervision
- StillFast: An End-to-End Approach for Short-Term Object Interaction Anticipation
- VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking
- Unmasked Teacher: Towards Training-Efficient Video Foundation Models
- Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
- Interaction Region Visual Transformer for Egocentric Action Anticipation
- VICRegL: Self-Supervised Learning of Local Visual Features
- OmniMAE: Single Model Masked Pretraining on Images and Videos
- Patch-level Representation Learning for Self-supervised Vision Transformers
- Self-Supervised Visual Representation Learning with Semantic Grouping
- Masked Autoencoders As Spatiotemporal Learners
- Self-Supervised Learning of Object Parts for Semantic Segmentation
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training
- Unsupervised Semantic Segmentation by Distilling Feature Correspondences
- Object discovery and representation networks
- data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
- MERLOT Reserve: Neural Script Knowledge through Vision and Language and Sound
- High-Resolution Image Synthesis with Latent Diffusion Models
- High-Resolution Image Synthesis with Latent Diffusion Models
- Masked Feature Prediction for Self-Supervised Visual Pre-Training
- SimMIM: A Simple Framework for Masked Image Modeling
- iBOT: Image BERT Pre-Training with Online Tokenizer
- Ego4D: Around the World in 3,000 Hours of Egocentric Video
- Unsupervised Object-Level Representation Learning from Scene Images
- BEiT: BERT Pre-Training of Image Transformers
- Segmenter: Transformer for Semantic Segmentation
- VICReg: Variance-Invariance-Covariance Regularization for\n Self-Supervised Learning
- Emerging Properties in Self-Supervised Vision Transformers
- Multiscale Vision Transformers
- An Empirical Study of Training Self-Supervised Vision Transformers
- ViViT: A Video Vision Transformer
- Vision Transformers for Dense Prediction
- Efficient Visual Pretraining with Contrastive Detection
- Is Space-Time Attention All You Need for Video Understanding?
- Betrayed by Motion: Camouflaged Object Discovery via Motion Segmentation
- Exploring Simple Siamese Representation Learning
- Exploring Simple Siamese Representation Learning
- Unsupervised Learning of Dense Visual Representations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Self-supervised Co-training for Video Representation Learning
- Denoising Diffusion Implicit Models
- Space-Time Correspondence as a Contrastive Random Walk
- Denoising Diffusion Probabilistic Models
- Unsupervised Learning of Visual Features by Contrasting Cluster Assignments
- Bootstrap your own latent: A new approach to self-supervised Learning
- A Simple Framework for Contrastive Learning of Visual Representations
- Self-Supervised Learning of Pretext-Invariant Representations
- Momentum Contrast for Unsupervised Visual Representation Learning
- Momentum Contrast for Unsupervised Visual Representation Learning
- Video Representation Learning by Dense Predictive Coding
- A Short Note on the Kinetics-700 Human Action Dataset
- Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-shot Cross-dataset Transfer
- Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-Shot Cross-Dataset Transfer
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips
- Data-Efficient Image Recognition with Contrastive Predictive Coding
- SlowFast Networks for Video Recognition
- YouTube-VOS: Sequence-to-Sequence Video Object Segmentation
- World Models
- Unsupervised Representation Learning by Predicting Image Rotations
- Rethinking Spatiotemporal Feature Learning: Speed-Accuracy Trade-offs in Video Classification
- A Closer Look at Spatiotemporal Convolutions for Action Recognition
- The "something something" video database for learning and evaluating visual common sense
- Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset
- The Kinetics Human Action Video Dataset
- The 2017 DAVIS Challenge on Video Object Segmentation
- Learning Features by Watching Objects Move
- RefineNet: Multi-Path Refinement Networks for High-Resolution Semantic Segmentation
- Context Encoders: Feature Learning by Inpainting
- The Cityscapes Dataset for Semantic Urban Scene Understanding
- Unsupervised Learning of Visual Representations by Solving Jigsaw Puzzles
- Colorful Image Colorization
- Unsupervised Visual Representation Learning by Context Prediction
- Unsupervised Visual Representation Learning by Context Prediction
- Learning image representations tied to ego-motion
- Learning to See by Moving
- Fast R-CNN
- Distilling the Knowledge in a Neural Network
- Long-term Recurrent Convolutional Networks for Visual Recognition and Description
- Vision meets robotics: The KITTI dataset
- LLaVA-Video: Video Instruction Tuning With Synthetic Data
- Guided Attention for Next Active Object @ EGO4D STA Challenge
- Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
- Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
Cited by
Discussions
Related