Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
2023/01/19 by Mahmoud Assran, Assran, Mahmoud, Quentin Duval +14 · 8 voices · 175 citations
Computer Science · #Domain Adaptation and Few-Shot Learning #Human Pose and Action Recognition #Advanced Neural Network Applications
paper · pdf · doi:10.48550/arxiv.2301.08243
Abstract
This paper demonstrates an approach for learning highly semantic image representations without relying on hand-crafted data-augmentations. We introduce the Image-based Joint-Embedding Predictive Architecture (I-JEPA), a non-generative approach for self-supervised learning from images. The idea behind I-JEPA is simple: from a single context block, predict the representations of various target blocks in the same image. A core design choice to guide I-JEPA towards producing semantic representations is the masking strategy; specifically, it is crucial to (a) sample target blocks with sufficiently large scale (semantic), and to (b) use a sufficiently informative (spatially distributed) context block. Empirically, when combined with Vision Transformers, we find I-JEPA to be highly scalable. For instance, we train a ViT-Huge/14 on ImageNet using 16 A100 GPUs in under 72 hours to achieve strong downstream performance across a wide range of tasks, from linear classification to object counting and depth prediction.
Cited by
- Music-JEPA: Learning a World Model of Sound from Action
- dRAE: Representation Autoencoder with Hyper-Spherical Codes
- IQ-JEPA: A Joint-Embedding Predictive Architecture with a Hermitian Vision Transformer for Sound Speed and Attenuation Estimation from Ultrasound IQ Data
- Hierarchical Self-Supervised Representation Learning Framework for Multivariate Time Series Grounded in ECG Analysis
- Depth-Regularized JEPA World Models Learn More Transferable Representations from Real Outdoor Robot Data
- LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition
- Toward Goal-Agnostic Joint-Embedding Predictive Control of Partial Differential Equations
- DWM: Separating World Effects from Actions in Latent World Models
- Multi-Horizon Consistency as Geometry: When Latent Dynamics Contract, and When They Do Not
- Applying JEPA-Style Predictive Learning to JA4-Derived Network Fingerprints
- Predictive Training with Latent Imagination for Visual Quadruped Navigation
- SAGE: Subgoal-Conditioned Action Generation for Latent World Model Planning
- Adaptive Compute in Latent World Models: When Depth Helps, Hurts, or Doesn't Matter
- REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion
- Joint-Embedding Predictive Architecture for Sensor-based Activity Recognition
- SC-JEPA: Stabilizing Latent Predictive Learning for Time-Series Anomaly Prediction
- DSWorld: A Data Science World Model for Efficient Autonomous Agents
- Orca: The World is in Your Mind
- MIDI-RAE-JEPA: Hierarchical Representation Learning and Generation for Symbolic Music
- AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning
- Learning Spatio-Temporal Foundation Models from Pure Synthetic Data
- A Compositional Framework for Open-ended Intelligence
- Information bottleneck for learning the phase space of dynamics from high-dimensional experimental data
- Why AI systems don't learn and what to do about it: Lessons on autonomous learning from cognitive science
- AI Must Embrace Specialization via Superhuman Adaptable Intelligence
- Learning to Detect Cross-Modal Negation: An Analysis of Latent Representations and an Attention-Based Solution
- Remapping and navigation of an embedding space via error minimization: a fundamental organizational principle of cognition in natural and artificial systems
- MuM: Multi-View Masked Image Modeling for 3D Vision
- Predicting upcoming visual features during eye movements yields scene representations aligned with human visual cortex
- What We Don't C: Manifold Disentanglement for Structured Discovery
- Understanding neural circuit principles for representation learning through joint-embedding predictive architectures
- Chimère Ω — blueprint for a physico-cognitively inspired local-first LLM runtime
- Beyond Sensor Data: Foundation Models of Behavioral Data from Wearables Improve Health Predictions
- AI4SNOW SnowGalileo Datasets and Model Checkpoints
- From Pixels to Components: Eigenvector Masking for Visual Representation Learning
- Web World Models
- The JEPA Paradox in Language: The Geometry of Linguistic Alternatives
- Normalizing Trajectory Models
- A satellite foundation model for improved wealth monitoring
- N0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens
- Artificial General Intelligence (AGI)-Native Wireless Systems: A Journey Beyond 6G
- The Semantic Least-Energy Principle: A Hypothesis for Intelligence
- FeelWorld: Visuo-Tactile World Model for Hierarchical Contact Prediction and Planning
- Medical world models in healthcare: foundations, applications, and challenges for trustworthy clinical translation
- 6G: From Connectivity Infrastructure to Guaranteed Digital Services
- How Much MRI Preprocessing Is Enough? A Cost-Utility Study for Brain MRI Foundation Models
- Autoregressive One-Step Generative Modeling for Dynamical System Forecasting
- The Cartesian Cut in Agentic AI
- Self-Distillation of Hidden Layers for Self-Supervised Representation Learning
- PhysFire-WM: A Physics-Informed World Model for Emulating Fire Spread Dynamics
- SLIM-Brain: A Data- and Training-Efficient Foundation Model for fMRI Data Analysis
- Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
- FGDCC: Fine-Grained Deep Cluster Categorization -- A Framework for Intra-Class Variability Problems in Plant Classification
- Zero-Shot Segmentation through Prototype-Guidance for Multi-Label Plant Species Identification
- JEPA-Reasoner: Decoupling Latent Reasoning from Token Generation
- Next-Embedding Prediction Makes Strong Vision Learners
- Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models
- In Pursuit of Pixel Supervision for Visual Pre-training
- Magnification-Aware Distillation (MAD): A Self-Supervised Framework for Unified Representation Learning in Gigapixel Whole-Slide Images
- Data-driven modelling of autonomous and forced dynamical systems
- VL-JEPA: Joint Embedding Predictive Architecture for Vision-language
- What matters for Representation Alignment: Global Information or Spatial Structure?
- Closing the Train-Test Gap in World Models for Gradient-Based Planning
- From SAM to DINOv2: Towards Distilling Foundation Models to Lightweight Baselines for Generalized Polyp Segmentation
- Is Generation Required for Data-Efficient Perception?
- JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention
- CLARITY: Medical World Model for Guiding Treatment Decisions by Modeling Context-Aware Disease Trajectories in Latent Space
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
- BiTAgent: A Task-Aware Modular Framework for Bidirectional Coupling between Multimodal Large Language Models and World Models
- Generalized Event Partonomy Inference with Structured Hierarchical Predictive Learning
- Cosine-Similarity Methods for Efficient Training and Sampling in High-Dimensional Latent Spaces
- CrossJEPA: Cross-Modal Joint-Embedding Predictive Architecture for Efficient 3D Representation Learning from 2D Images
- stable-pretraining-v1: Foundation Model Research Made Simple
- DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture
- Align & Invert: Solving Inverse Problems with Diffusion and Flow-based Models via Representation Alignment
- POMA-3D: The Point Map Way to 3D Scene Understanding
- X-WIN: Building Chest Radiograph World Model via Predictive Sensing
- OlmoEarth: Stable Latent Image Modeling for Multimodal Earth Observation
- Do Blind Spots Matter for Word-Referent Mapping? A Computational Study with Infant Egocentric Video
- PriVi: Towards A General-Purpose Video Model For Primate Behavior In The Wild
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
- Leveraging unlabelled data for generalizable neural population decoding
- Mutual information and task-relevant latent dimensionality
- FlowFeat: Pixel-Dense Embedding of Motion Profiles
- SAMora: Enhancing SAM through Hierarchical Self-Supervised Pre-Training for Medical Images
- Cambrian-S: Towards Spatial Supersensing in Video
- ProM3E: Probabilistic Masked MultiModal Embedding Model for Ecology
- Graph-Enhanced Policy Optimization in LLM Agent Training
- Controlling Contrastive Self-Supervised Learning with Knowledge-Driven Multiple Hypothesis: Application to Beat Tracking
- Enfold: Folding World-Generator Computation into Predictive Representations for Efficient Embodied Control
- JEPADepth: Masked Predictive Representation Learning for Self-Supervised Monocular Depth Estimation
- Rad-JEPA 3D: Radiology Joint-Embedding Predictive Model for 3D Computed Tomography
- Analyzing Image Encoder Choices and Graph Homophily in GCN Frameworks for Breast Ultrasound Classification
- Face-Trace: Open-Set Attribution and Progressive Discovery of Synthetic Face Generators
- Temporally Centered SIGReg Improves Multi-Task LeWorldModel Learning: From Analysis to Method
- TweetyBERT: Automated parsing of birdsong through self-supervised machine learning
- Seeking the Unfamiliar but Memorable: Conceptual Creativity as Meta-Learning
- NeuroAI and Beyond: Bridging Between Advances in Neuroscience and ArtificialIntelligence
- Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models
- Eigenfunction Extraction for Ordered Representation Learning
- Perception Learning: A Formal Separation of Sensory Representation Learning from Decision Learning
- Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations
- ROPES: Robotic Pose Estimation via Score-Based Causal Representation Learning
- H-SPLID: HSIC-based Saliency Preserving Latent Information Decomposition
- Why Prototypes Collapse: Diagnosing and Preventing Partial Collapse in Prototypical Self-Supervised Learning
- MoAlign: Motion-Centric Representation Alignment for Video Diffusion Models
- Exploring Structural Degradation in Dense Representations for Self-supervised Learning
- Terra: Explorable Native 3D World Model with Point Latents
- Semantic representations emerge in biologically inspired ensembles of cross-supervising neural networks
- Scaling Vision Transformers for Functional MRI with Flat Maps
- Unifying Vision-Language Latents for Zero-label Image Caption Enhancement
- Diffusion Transformers with Representation Autoencoders
- Learning at the Speed of Physics: Equilibrium Propagation on Oscillator Ising Machines
- Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
- Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight
- VA-Adapter: Adapting Ultrasound Foundation Model to Echocardiography Probe Guidance
- Gaussian Embeddings: How JEPAs Secretly Learn Your Data Density
- Midway Network: Learning Representations for Recognition and Motion from Latent Dynamics
- Foundation Visual Encoders Are Secretly Few-Shot Anomaly Detectors
- Using predefined vector systems as latent space configuration for neural network supervised training on data with arbitrarily large number of classes
- Unsupervised Transformer Pre-Training for Images: Self-Distillation, Mean Teachers, and Random Crops
- A Hybrid Co-Finetuning Approach for Visual Bug Detection in Video Games
- Fusing Multi- and Hyperspectral Satellite Data for Harmful Algal Bloom Monitoring with Self-Supervised and Hierarchical Deep Learning
- JEPA-T: Joint-Embedding Predictive Architecture with Text Fusion for Image Generation
- Joint Embeddings Go Temporal
- Benchmarking ECG Foundational Models: A Reality Check Across Clinical Tasks
- Event-based Facial Keypoint Alignment via Cross-Modal Fusion Attention and Self-Supervised Multi-Event Representation Learning
- Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers
- Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification
- Fast Feature Field (F3): A Predictive Representation of Events
- Brain Harmony: A Multimodal Foundation Model Unifying Morphology and Function into 1D Tokens
- GenView++: Unifying Adaptive Generative Augmentation and Quality-Driven Supervision for Contrastive Representation Learning
- WavJEPA: Semantic learning unlocks robust audio foundation models for raw waveforms
- Orochi: Versatile Biomedical Image Processor
- Learning the Neighborhood: Contrast-Free Multimodal Self-Supervised Molecular Graph Pretraining
- SSVIF: Self-Supervised Segmentation-Oriented Visible and Infrared Image Fusion
- Embodied AI: From LLMs to World Models
- MeshFM: 2D Features Are All You Need for 3D Shape Understanding
- SiamJEPA: On the Role of Siamese Student Encoders in JEPA
- Large scale cross-regional remote sensing flood monitoring framework for operative mapping and impact analysis
- ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow
- One Future, Every Robot: Label-Efficient Collective-State Prediction with Decentralized JEPA
- What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations
- ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution
- Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
- Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis
- The Platonic Universe: Do Foundation Models See the Same Sky?
- Knowledge Transfer from Interaction Learning
- Can multimodal representation learning by alignment preserve modality-specific information?
- Enhancing Semantic Segmentation with Continual Self-Supervised Pre-training
- Remote Sensing-Oriented World Model
- UNIV: Unified Foundation Model for Infrared and Visible Modalities
- Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
- BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition
- RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
- Self-supervised learning of imaging and clinical signatures using a multimodal joint-embedding predictive architecture
- Masked Feature Modeling Enhances Adaptive Segmentation
- MMMS: Multi-Modal Multi-Surface Interactive Segmentation
- Label-Efficient Grasp Joint Prediction with Point-JEPA
- Building a General SimCLR Self-Supervised Foundation Model Across Neurological Diseases to Advance 3D Brain MRI Diagnoses
- BenchECG and xECG: a benchmark and baseline for ECG foundation models
- Why and How Auxiliary Tasks Improve JEPA Representations
- LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
- Diffusion-Based Action Recognition Generalizes to Untrained Domains
- Planning with Reasoning using Vision Language World Model
- Maybe you don't need a U-Net: convolutional feature upsampling for materials micrograph segmentation
- Generalizable Object Re-Identification via Visual In-Context Prompting
- From Linearity to Non-Linearity: How Masked Autoencoders Capture Spatial Correlations
- EmoSLLM: Parameter-Efficient Adaptation of LLMs for Speech Emotion Recognition
- Contrastive Representations for Temporal Reasoning
- MATPAC++: Enhanced Masked Latent Prediction for Self-Supervised Audio Representation Learning
- MAESTRO: Masked AutoEncoders for Multimodal, Multitemporal, and Multispectral Earth Observation Data
- Learning State-Space Models of Dynamic Systems from Arbitrary Data using Joint Embedding Predictive Architectures
- GoViG: Goal-Conditioned Visual Navigation Instruction Generation
- SegDAC: Visual Generalization in Reinforcement Learning via Dynamic Object Tokens
Discussions
- We think cortex might function like a JEPA. It looks like prediction errors in layer 2/3 are not computed against input (as is the idea in predictive processing), but against a representation in laten [bsky, 45 points, 2 comments]
- Self-Supervised Learning from Images with JEPA (2023) [hn, 40 points, 10 comments]
- This seems to be the original paper: arxiv.org/abs/2301.08243 (I’ve coauthored with Mike Rabbat which makes me have a collaboration distance of 2 from LeCun, that’s pretty sweet) [bsky, 3 points, 1 comments]
- Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture [lemmy, 3 points, 1 comments]
- Self-Supervised Learning from Images with JEPA https://arxiv.org/abs/2301.08243 (https://news.ycombinator.com/item?id=43512657) [bsky, 1 points, 0 comments]
- Self-Supervised Learning from Images with a Joint-Embedding Predictive Archi [hn, 1 points, 0 comments]
- Self-Supervised Learning from Images with JEPA https://arxiv.org/abs/2301.08243 [bsky, 1 points, 0 comments]
- Self-Supervised Learning from Images with JEPA #HackerNews https://arxiv.org/abs/2301.08243 [bsky, 0 points, 0 comments]
Related