Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
2023/01/19 by Mahmoud Assran, Assran, Mahmoud, Quentin Duval +14 · 8 voices · 284 citations
Computer Science · Mathematics · #Advanced Neural Network Applications #Architecture #Artificial intelligence #Artificial neural network #Block (permutation group theory) #Computer science #Context (archaeology) #Domain Adaptation and Few-Shot Learning #Embedding #Feature learning #Human Pose and Action Recognition #Joint (building) #Machine learning #Mathematics #Pattern recognition (psychology) #Scalability #Supervised learning #Transformer
paper · pdf · doi:10.48550/arxiv.2301.08243
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/01/19 · openalex created_date 2023/01/21 · openalex updated_date 2026/07/28
Abstract
This paper demonstrates an approach for learning highly semantic image representations without relying on hand-crafted data-augmentations. We introduce the Image-based Joint-Embedding Predictive Architecture (I-JEPA), a non-generative approach for self-supervised learning from images. The idea behind I-JEPA is simple: from a single context block, predict the representations of various target blocks in the same image. A core design choice to guide I-JEPA towards producing semantic representations is the masking strategy; specifically, it is crucial to (a) sample target blocks with sufficiently large scale (semantic), and to (b) use a sufficiently informative (spatially distributed) context block. Empirically, when combined with Vision Transformers, we find I-JEPA to be highly scalable. For instance, we train a ViT-Huge/14 on ImageNet using 16 A100 GPUs in under 72 hours to achieve strong downstream performance across a wide range of tasks, from linear classification to object counting and depth prediction.
Cited by
- Music-JEPA: Learning a World Model of Sound from Action
- dRAE: Representation Autoencoder with Hyper-Spherical Codes
- IQ-JEPA: A Joint-Embedding Predictive Architecture with a Hermitian Vision Transformer for Sound Speed and Attenuation Estimation from Ultrasound IQ Data
- Hierarchical Self-Supervised Representation Learning Framework for Multivariate Time Series Grounded in ECG Analysis
- Depth-Regularized JEPA World Models Learn More Transferable Representations from Real Outdoor Robot Data
- LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition
- Toward Goal-Agnostic Joint-Embedding Predictive Control of Partial Differential Equations
- DWM: Separating World Effects from Actions in Latent World Models
- Multi-Horizon Consistency as Geometry: When Latent Dynamics Contract, and When They Do Not
- Applying JEPA-Style Predictive Learning to JA4-Derived Network Fingerprints
- Predictive Training with Latent Imagination for Visual Quadruped Navigation
- SAGE: Subgoal-Conditioned Action Generation for Latent World Model Planning
- Adaptive Compute in Latent World Models: When Depth Helps, Hurts, or Doesn't Matter
- REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion
- Joint-Embedding Predictive Architecture for Sensor-based Activity Recognition
- SC-JEPA: Stabilizing Latent Predictive Learning for Time-Series Anomaly Prediction
- DSWorld: A Data Science World Model for Efficient Autonomous Agents
- Orca: The World is in Your Mind
- MIDI-RAE-JEPA: Hierarchical Representation Learning and Generation for Symbolic Music
- AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning
- Learning Spatio-Temporal Foundation Models from Pure Synthetic Data
- A Compositional Framework for Open-ended Intelligence
- Information bottleneck for learning the phase space of dynamics from high-dimensional experimental data
- Why AI systems don't learn and what to do about it: Lessons on autonomous learning from cognitive science
- AI Must Embrace Specialization via Superhuman Adaptable Intelligence
- Learning to Detect Cross-Modal Negation: An Analysis of Latent Representations and an Attention-Based Solution
- Remapping and navigation of an embedding space via error minimization: a fundamental organizational principle of cognition in natural and artificial systems
- MuM: Multi-View Masked Image Modeling for 3D Vision
- Predicting upcoming visual features during eye movements yields scene representations aligned with human visual cortex
- What We Don't C: Manifold Disentanglement for Structured Discovery
- Understanding neural circuit principles for representation learning through joint-embedding predictive architectures
- Chimère Ω — blueprint for a physico-cognitively inspired local-first LLM runtime
- Beyond Sensor Data: Foundation Models of Behavioral Data from Wearables Improve Health Predictions
- AI4SNOW SnowGalileo Datasets and Model Checkpoints
- From Pixels to Components: Eigenvector Masking for Visual Representation Learning
- Web World Models
- The JEPA Paradox in Language: The Geometry of Linguistic Alternatives
- Normalizing Trajectory Models
- A satellite foundation model for improved wealth monitoring
- N0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens
- Artificial General Intelligence (AGI)-Native Wireless Systems: A Journey Beyond 6G
- The Semantic Least-Energy Principle: A Hypothesis for Intelligence
- FeelWorld: Visuo-Tactile World Model for Hierarchical Contact Prediction and Planning
- Medical world models in healthcare: foundations, applications, and challenges for trustworthy clinical translation
- 6G: From Connectivity Infrastructure to Guaranteed Digital Services
- How Much MRI Preprocessing Is Enough? A Cost-Utility Study for Brain MRI Foundation Models
- Autoregressive One-Step Generative Modeling for Dynamical System Forecasting
- The Cartesian Cut in Agentic AI
- Self-Distillation of Hidden Layers for Self-Supervised Representation Learning
- PhysFire-WM: A Physics-Informed World Model for Emulating Fire Spread Dynamics
- SLIM-Brain: A Data- and Training-Efficient Foundation Model for fMRI Data Analysis
- Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
- FGDCC: Fine-Grained Deep Cluster Categorization -- A Framework for Intra-Class Variability Problems in Plant Classification
- Zero-Shot Segmentation through Prototype-Guidance for Multi-Label Plant Species Identification
- JEPA-Reasoner: Decoupling Latent Reasoning from Token Generation
- Next-Embedding Prediction Makes Strong Vision Learners
- Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models
- In Pursuit of Pixel Supervision for Visual Pre-training
- Magnification-Aware Distillation (MAD): A Self-Supervised Framework for Unified Representation Learning in Gigapixel Whole-Slide Images
- Data-driven modelling of autonomous and forced dynamical systems
- VL-JEPA: Joint Embedding Predictive Architecture for Vision-language
- What matters for Representation Alignment: Global Information or Spatial Structure?
- Closing the Train-Test Gap in World Models for Gradient-Based Planning
- From SAM to DINOv2: Towards Distilling Foundation Models to Lightweight Baselines for Generalized Polyp Segmentation
- Is Generation Required for Data-Efficient Perception?
- JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention
- CLARITY: Medical World Model for Guiding Treatment Decisions by Modeling Context-Aware Disease Trajectories in Latent Space
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
- BiTAgent: A Task-Aware Modular Framework for Bidirectional Coupling between Multimodal Large Language Models and World Models
- Generalized Event Partonomy Inference with Structured Hierarchical Predictive Learning
- Cosine-Similarity Methods for Efficient Training and Sampling in High-Dimensional Latent Spaces
- CrossJEPA: Cross-Modal Joint-Embedding Predictive Architecture for Efficient 3D Representation Learning from 2D Images
- stable-pretraining-v1: Foundation Model Research Made Simple
- DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture
- Align & Invert: Solving Inverse Problems with Diffusion and Flow-based Models via Representation Alignment
- POMA-3D: The Point Map Way to 3D Scene Understanding
- X-WIN: Building Chest Radiograph World Model via Predictive Sensing
- OlmoEarth: Stable Latent Image Modeling for Multimodal Earth Observation
- Do Blind Spots Matter for Word-Referent Mapping? A Computational Study with Infant Egocentric Video
- PriVi: Towards A General-Purpose Video Model For Primate Behavior In The Wild
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
- Leveraging unlabelled data for generalizable neural population decoding
- Mutual information and task-relevant latent dimensionality
- FlowFeat: Pixel-Dense Embedding of Motion Profiles
- SAMora: Enhancing SAM through Hierarchical Self-Supervised Pre-Training for Medical Images
- Cambrian-S: Towards Spatial Supersensing in Video
- ProM3E: Probabilistic Masked MultiModal Embedding Model for Ecology
- Graph-Enhanced Policy Optimization in LLM Agent Training
- Controlling Contrastive Self-Supervised Learning with Knowledge-Driven Multiple Hypothesis: Application to Beat Tracking
- Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control
- JEPADepth: Masked Predictive Representation Learning for Self-Supervised Monocular Depth Estimation
- Rad-JEPA 3D: Radiology Joint-Embedding Predictive Model for 3D Computed Tomography
- Analyzing Image Encoder Choices and Graph Homophily in GCN Frameworks for Breast Ultrasound Classification
- Face-Trace: Open-Set Attribution and Progressive Discovery of Synthetic Face Generators
- Temporally Centered SIGReg Improves Multi-Task LeWorldModel Learning: From Analysis to Method
- TweetyBERT: Automated parsing of birdsong through self-supervised machine learning
- Seeking the Unfamiliar but Memorable: Conceptual Creativity as Meta-Learning
- A quantum semantic framework for natural language processing
- NeuroAI and Beyond: Bridging Between Advances in Neuroscience and ArtificialIntelligence
- Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models
- Eigenfunction Extraction for Ordered Representation Learning
- Perception Learning: A Formal Separation of Sensory Representation Learning from Decision Learning
- Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations
- ROPES: Robotic Pose Estimation via Score-Based Causal Representation Learning
- H-SPLID: HSIC-based Saliency Preserving Latent Information Decomposition
- Why Prototypes Collapse: Diagnosing and Preventing Partial Collapse in Prototypical Self-Supervised Learning
- MoAlign: Motion-Centric Representation Alignment for Video Diffusion Models
- MSSDF: Modality-Shared Self-supervised Distillation for High-Resolution Multi-modal Remote Sensing Image Learning
- Exploring Structural Degradation in Dense Representations for Self-supervised Learning
- Terra: Explorable Native 3D World Model with Point Latents
- Semantic representations emerge in biologically inspired ensembles of cross-supervising neural networks
- Scaling Vision Transformers for Functional MRI with Flat Maps
- Unifying Vision-Language Latents for Zero-label Image Caption Enhancement
- Diffusion Transformers with Representation Autoencoders
- Learning at the Speed of Physics: Equilibrium Propagation on Oscillator Ising Machines
- Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
- Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight
- VA-Adapter: Adapting Ultrasound Foundation Model to Echocardiography Probe Guidance
- Gaussian Embeddings: How JEPAs Secretly Learn Your Data Density
- Midway Network: Learning Representations for Recognition and Motion from Latent Dynamics
- Foundation Visual Encoders Are Secretly Few-Shot Anomaly Detectors
- Using predefined vector systems as latent space configuration for neural network supervised training on data with arbitrarily large number of classes
- Unsupervised Transformer Pre-Training for Images: Self-Distillation, Mean Teachers, and Random Crops
- A Hybrid Co-Finetuning Approach for Visual Bug Detection in Video Games
- Fusing Multi- and Hyperspectral Satellite Data for Harmful Algal Bloom Monitoring with Self-Supervised and Hierarchical Deep Learning
- JEPA-T: Joint-Embedding Predictive Architecture with Text Fusion for Image Generation
- Joint Embeddings Go Temporal
- Benchmarking ECG FMs: A Reality Check Across Clinical Tasks
- Event-based Facial Keypoint Alignment via Cross-Modal Fusion Attention and Self-Supervised Multi-Event Representation Learning
- Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers
- Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification
- Fast Feature Field (F3): A Predictive Representation of Events
- Brain Harmony: A Multimodal Foundation Model Unifying Morphology and Function into 1D Tokens
- GenView++: Unifying Adaptive Generative Augmentation and Quality-Driven Supervision for Contrastive Representation Learning
- WavJEPA: Semantic learning unlocks robust audio foundation models for raw waveforms
- Orochi: Versatile Biomedical Image Processor
- Learning the Neighborhood: Contrast-Free Multimodal Self-Supervised Molecular Graph Pretraining
- SSVIF: Self-Supervised Segmentation-Oriented Visible and Infrared Image Fusion
- Embodied AI: From LLMs to World Models
- MeshFM: 2D Features Are All You Need for 3D Shape Understanding
- SiamJEPA: On the Role of Siamese Student Encoders in JEPA
- Large scale cross-regional remote sensing flood monitoring framework for operative mapping and impact analysis
- ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow
- One Future, Every Robot: Label-Efficient Collective-State Prediction with Decentralized JEPA
- What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations
- ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution
- Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
- Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis
- The Platonic Universe: Do Foundation Models See the Same Sky?
- Knowledge Transfer from Interaction Learning
- Can multimodal representation learning by alignment preserve modality-specific information?
- Enhancing Semantic Segmentation with Continual Self-Supervised Pre-training
- Remote Sensing-Oriented World Model
- UNIV: Unified Foundation Model for Infrared and Visible Modalities
- Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
- BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition
- RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
- Self-supervised learning of imaging and clinical signatures using a multimodal joint-embedding predictive architecture
- Masked Feature Modeling Enhances Adaptive Segmentation
- MMMS: Multi-Modal Multi-Surface Interactive Segmentation
- Label-Efficient Grasp Joint Prediction with Point-JEPA
- Building a General SimCLR Self-Supervised Foundation Model Across Neurological Diseases to Advance 3D Brain MRI Diagnoses
- BenchECG and xECG: a benchmark and baseline for ECG foundation models
- Why and How Auxiliary Tasks Improve JEPA Representations
- LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
- Diffusion-Based Action Recognition Generalizes to Untrained Domains
- Planning with Reasoning using Vision Language World Model
- Maybe you don't need a U-Net: convolutional feature upsampling for materials micrograph segmentation
- Generalizable Object Re-Identification via Visual In-Context Prompting
- From Linearity to Non-Linearity: How Masked Autoencoders Capture Spatial Correlations
- EmoSLLM: Parameter-Efficient Adaptation of LLMs for Speech Emotion Recognition
- Contrastive Representations for Temporal Reasoning
- MATPAC++: Enhanced Masked Latent Prediction for Self-Supervised Audio Representation Learning
- MAESTRO: Masked AutoEncoders for Multimodal, Multitemporal, and Multispectral Earth Observation Data
- Learning State-Space Models of Dynamic Systems from Arbitrary Data using Joint Embedding Predictive Architectures
- GoViG: Goal-Conditioned Visual Navigation Instruction Generation
- SegDAC: Visual Generalization in Reinforcement Learning via Dynamic Object Tokens
- ReSim: Reliable World Simulation for Autonomous Driving
- SpecBPP: A Self-Supervised Learning Approach for Hyperspectral Representation and Soil Organic Carbon Estimation
- eMargin: Revisiting Contrastive Learning with Margin-Based Separation
- Unlocking the Working Memory of Large Language Models for Latent Reasoning
- PRAGMA: Revolut Foundation Model
- Masked Depth Modeling for Spatial Perception
- Critique of impure reason: Unveiling the reasoning behaviour of medical large language models
- SV3.3B: A Sports Video Understanding Model for Action Recognition
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- MoSiC: Optimal-Transport Motion Trajectory for Dense Self-Supervised Learning
- Foundation Models in Medical Imaging: A Review and Outlook
- Task Priors: Enhancing Model Evaluation by Considering the Entire Space of Downstream Tasks
- Scalable Spatiotemporal Inference with Biased Scan Attention Transformer Neural Processes
- AquaJEPA: Action-Conditioned Multimodal Predictive Representations for Underwater Robot Dynamics
- GreenHyperSpectra: A multi-source hyperspectral dataset for global vegetation trait prediction
- Tractable Representation Learning with Probabilistic Circuits
- Physics-Aligned Self-Supervised Learning for Scientific Imaging
- From Video to EEG: Adapting Joint Embedding Predictive Architecture to Uncover Saptiotemporal Dynamics in Brain Signal Analysis
- Is Visual in-Context Learning for Compositional Medical Tasks within Reach?
- A Survey: Learning Embodied Intelligence from Physical Simulators and World Models
- Resilient-Native and Intelligent Next-Generation Wireless Systems: Key Enablers, Foundations, and Applications
- Active Inference AI Systems for Scientific Discovery
- Boosting Generative Adversarial Transferability with Self-supervised Vision Transformer Features
- Asymmetric Dual Self-Distillation for 3D Self-Supervised Representation Learning
- Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
- SMARTIES: Spectrum-Aware Multi-Sensor Auto-Encoder for Remote Sensing Images
- BrainSymphony: A Transformer-Driven Fusion of fMRI Time Series and Structural Connectivity
- Joint Embedding Predictive Architecture for self-supervised pretraining on polymer molecular graphs
- Enhancing VICReg: Random-Walk Pairing for Improved Generalization and Better Global Semantics Capturing
- Time-Contrastive Pretraining for In-Context Image and Video Segmentation
- SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding
- A Survey of State Representation Learning for Deep Reinforcement Learning
- Dual Perspectives on Non-Contrastive Self-Supervised Learning
- HiT-JEPA: A Hierarchical Self-supervised Trajectory Embedding Framework for Similarity Computation
- Discrete JEPA: Learning Discrete Token Representations without Reconstruction
- Visual Pre-Training on Unlabeled Images using Reinforcement Learning
- Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models
- WorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural Planning
- Object-level Self-Distillation for Vision Pretraining
- How PARTs assemble into wholes: Learning the relative composition of images
- ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads
- seq-JEPA: Autoregressive Predictive Learning of Invariant-Equivariant World Models
- Latent Reasoning via Sentence Embedding Prediction
- Object Concepts Emerge from Motion
- Vision Transformers with Self-Distilled Registers
- Multimodal Federated Learning With Missing Modalities through Feature Imputation Network
- OSVI-WM: One-Shot Visual Imitation for Unseen Tasks using World-Model-Guided Trajectory Generation
- Self-Organizing Visual Prototypes for Non-Parametric Representation Learning
- LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-Supervision
- SpikeGen: Decoupled "Rods and Cones" Visual Representation Processing with Latent Generative Framework
- Structured Proxy Features for Multimodal NSCLC Survival Prediction from Pretreatment CT
- Bootstrapping your behavior: a new pretraining strategy for user behavior sequence data
- Time to Embed: Unlocking Foundation Models for Time Series with Channel Descriptions
- Collaborative Unlabeled Data Optimization
- Joint Embedding vs Reconstruction: Provable Benefits of Latent Space Prediction for Self Supervised Learning
- Asleep at the Wheel: JEPA's Limitations in Evaluating Novel Driving Data
- Surgical Foundation Model Leveraging Compression and Entropy Maximization for Image-Guided Surgical Assistance
- Self-supervised perception for tactile skin covered dexterous hands
- UniVLA: Learning to Act Anywhere with Task-centric Latent Actions
- Towards a Unified Representation Evaluation Framework Beyond Downstream Tasks
- How to Train an Oscillator Ising Machine using Equilibrium Propagation
- Contextures: Representations from Contexts
- Quo Vadis, World Modeling?
- VLAFlow: A Unified Training Framework for Vision-Language-Action Models via Co-training and Future Latent Alignment
- A Lightweight Library for Energy-Based Joint-Embedding Predictive Architectures
- V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
- Fast LeWorldModel
- JEDEL: Zero-Shot DNA-Encoded Library Design for Early-Stage Drug Discovery
- When to Align, When to Predict: A Phase Diagram for Multimodal Learning
- Next Embedding Prediction Makes World Models Stronger
- World Action Models are Zero-shot Policies
- On the Dynamics of Observation and Semantics
- Distilling LLM Feedback for Lean Theorem Proving
- When Does LeJEPA Learn a World Model?
- SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer
- On Training in Imagination
- You Don't Need Strong Assumptions: Visual Representation Learning via Temporal Differences
- WEAVER, Better, Faster, Longer: An Effective World Model for Robotic Manipulation
- AI+HW 2035: Shaping the Next Decade
- Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation
- CARL: Camera-Agnostic Representation Learning for Spectral Image Analysis
- SemanticMoments: Training-Free Motion Similarity via Third Moment Features
- FOUND-AF: Benchmarking ECG Foundation Models for Atrial Fibrillation Detection
- Accountability Asymmetry and Structural Trust in Autonomous AI Systems
- A Genealogy of Foundation Models in Remote Sensing
- Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation
- SJEPA: Learning Elegant Latent Dynamics with Hybrid Symbolic-Neural Predictors
- JEPA for RL: Investigating Joint-Embedding Predictive Architectures for Reinforcement Learning
- SparseJEPA: Sparse Representation Learning of Joint Embedding Predictive Architectures
- Locate 3D: Real-World Object Localization via Self-Supervised Learning in 3D
- CheXWorld: Exploring Image World Modeling for Radiograph Representation Learning
- Enhanced Pruning Strategy for Multi-Component Neural Architectures Using Component-Aware Graph Analysis
- Can Masked Autoencoders Also Listen to Birds?
- EchoWorld: Learning Motion-Aware World Models for Echocardiography Probe Guidance
- H3GNNs: Harmonizing Heterophily and Homophily in GNNs via Joint Structural Node Encoding and Self-Supervised Learning
- SR-JEPA: Learning Predictive Latent State in 3D Scenes
- PhyLatent: Learning Dynamics-Relevant Representations for JEPA World Models
- BioM-JEPA: joint-embedding prediction of graph-connected gene blocks in single cells
- Bar-JEPA: Extracting Values from Bar Chart with Joint-Embedding Predictive Architecture
- Discrete energy as an exact label-free training objective for finite-element surrogates
- REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion Transformers
- Unveiling Contrastive Learning's Capability of Neighborhood Aggregation for Collaborative Filtering
- Evolved Hierarchical Masking for Self-Supervised Learning
- Enhancing knowledge retention for continual learning with domain-specific adapters and features gating
- JEPA4Rec: Learning Effective Language Representations for Sequential Recommendation via Joint Embedding Predictive Architecture
- Masked Scene Modeling: Narrowing the Gap Between Supervised and Self-Supervised Learning in 3D Scene Understanding
- REJEPA: A Novel Joint-Embedding Predictive Architecture for Efficient Remote Sensing Image Retrieval
Discussions
- We think cortex might function like a JEPA. It looks like prediction errors in layer 2/3 are not computed against input (as is the idea in predictive processing), but against a representation in laten [bsky, 45 points, 2 comments]
- Self-Supervised Learning from Images with JEPA (2023) [hn, 40 points, 10 comments]
- This seems to be the original paper: arxiv.org/abs/2301.08243 (I’ve coauthored with Mike Rabbat which makes me have a collaboration distance of 2 from LeCun, that’s pretty sweet) [bsky, 3 points, 1 comments]
- Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture [lemmy, 3 points, 1 comments]
- Self-Supervised Learning from Images with JEPA https://arxiv.org/abs/2301.08243 (https://news.ycombinator.com/item?id=43512657) [bsky, 1 points, 0 comments]
- Self-Supervised Learning from Images with a Joint-Embedding Predictive Archi [hn, 1 points, 0 comments]
- Self-Supervised Learning from Images with JEPA https://arxiv.org/abs/2301.08243 [bsky, 1 points, 0 comments]
- Self-Supervised Learning from Images with JEPA #HackerNews https://arxiv.org/abs/2301.08243 [bsky, 0 points, 0 comments]
Related