LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
2026/03/13 by Lucas Maes, Quentin Le Lidec, Damien Scieur +2 · 14 voices · 5 citations
#cs.LG #cs.AI
paper · pdf
Abstract
Joint Embedding Predictive Architectures (JEPAs) offer a compelling framework for learning world models in compact latent spaces, yet existing methods remain fragile, relying on complex multi-term losses, exponential moving averages, pre-trained encoders, or auxiliary supervision to avoid representation collapse. In this work, we introduce LeWorldModel (LeWM), the first JEPA that trains stably end-to-end from raw pixels using only two loss terms: a next-embedding prediction loss and a regularizer enforcing Gaussian-distributed latent embeddings. This reduces tunable loss hyperparameters from six to one compared to the only existing end-to-end alternative. With ~15M parameters trainable on a single GPU in a few hours, LeWM plans up to 48x faster than foundation-model-based world models while remaining competitive across diverse 2D and 3D control tasks. Beyond control, we show that LeWM's latent space encodes meaningful physical structure through probing of physical quantities. Surprise evaluation confirms that the model reliably detects physically implausible events.
Citations
Cited by
Discussions
- LeWorldModel: Stable End-to-End Predictive Architecture from Pixels [hn, 6 points, 0 comments]
- Stable End-to-End Joint-Embedding Predictive Architecture from Pixels [hn, 5 points, 0 comments]
- LeWorldModel with Yann LeCun [hn, 2 points, 1 comments]
- Aquí lo tienes arxiv.org/abs/2603.19312 [bsky, 2 points, 2 comments]
- LeWorldModel: Stable E2E Joint-Embedding Predictive Architecture from Pixels [hn, 2 points, 0 comments]
- New breakthrough in World Models allows the AI to actually learn the underlying physics to make vision predictions with a tiny 15M param model, trainable on a single GPU arxiv.org/pdf/2603.19312 [bsky, 1 points, 1 comments]
- Paper review: LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels https://arxiv.org/pdf/2603.19312 Nice clean github: https://github.com/lucas-maes/le-wm (1/14) [bsky, 0 points, 1 comments]
- arxiv.org/abs/2603.19312 [bsky, 0 points, 0 comments]
- [2603.19312v1] LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels — A clear look at LeWorldModel, an end-to-end joint-embedding predictive setup that learns a stable w [bsky, 0 points, 0 comments]
- 今日読んだ論文 LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels arxiv.org/abs/2603.19312 JEPA的なやり方で表現空間における次状態予測損失で世界モデルを学習する。stop-gradientやEMAなし、次状態予測+SIGRegだけで学習に成功しているのは [bsky, 0 points, 0 comments]
- LeWorldModel strengthens Yann LeCun’s long standing argument that latent predictive world models may be far more efficient than brute force generative approaches for learning actionable structure from [bsky, 0 points, 0 comments]
- 30:26 - Isaac Ward — LeWorldModeling (https://arxiv.org/abs/2603.19312) 43:54 - Akshay Vegesna (@akshayvegesna) — Deep Learning is Not So Mysterious or Different (https://arxiv.org/abs/2503.02113) 51: [bsky, 0 points, 1 comments]
- Interesting references: arxiv.org/abs/2603.19312 [bsky, 0 points, 1 comments]
- ちなみにこれ。 arxiv.org/abs/2603.19312 [bsky, 0 points, 0 comments]
Related