JEPA (Joint-Embedding Predictive Architecture) is Yann LeCun’s proposed basis for world models. A JEPA predicts a representation of a future observation from a representation of a past observation rather than predicting raw pixels or tokens. The learned latent space lets the predictor ignore unpredictable surface details such as pixel noise, lighting changes, and paraphrase variation. It can then focus on the structural and semantic parts of the signal that support planning.
This section presents the JEPA variants in the order in which they were proposed.
JEPA (theory, 2022)
LeCun’s framing paper argues that generative prediction in observation space is the wrong objective for learning a world model. Such a model spends capacity on irrelevant pixel-level detail. LeCun proposes that perception, prediction, planning, and short- and long-term memory should instead operate on learned representations. Each later JEPA implements part of this design.- Paper: Yann LeCun (2022). A Path Towards Autonomous Machine Intelligence, OpenReview position paper.
- Companion lecture notes: Anna Dawid & Yann LeCun (2023). Introduction to Latent Variable Energy-Based Models: A Path Towards Autonomous Machine Intelligence, Les Houches Summer School notes on the energy-based formulation behind JEPA.
I-JEPA (images)
I-JEPA is the first concrete JEPA. Given a single image, it masks target blocks. A predictor operates on encoded patches rather than pixels and predicts the representations of the masked blocks from the unmasked context. I-JEPA matches masked-autoencoder quality on ImageNet without a pixel-level reconstruction loss. This result supports latent-space prediction for static images.- Paper: Assran et al. (2023). Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture, CVPR 2023.
MC-JEPA (motion + content)
MC-JEPA extends I-JEPA beyond a single still image. It jointly learns content (what is in the scene) and motion (how it flows) from video pairs. A shared encoder produces representations for semantic tasks such as classification and segmentation, as well as for optical-flow estimation. The results show that the JEPA objective can learn temporal structure while retaining static-image quality.- Paper: Bardes, Ponce, LeCun (2023). MC-JEPA: A Joint-Embedding Predictive Architecture for Self-Supervised Learning of Motion and Content Features, arXiv.
V-JEPA (video / dynamics)
V-JEPA applies latent-space prediction to full video. Its encoder processes many seconds of video, and its predictor reconstructs the representations of masked space-time regions. The resulting features transfer to downstream video-understanding tasks without training with a generative pixel loss. V-JEPA 2 (2025) scales the model to billions of parameters and serves as the visual backbone for several later VLA and VL-JEPA systems.- Paper: Bardes et al. (2024). Revisiting Feature Prediction for Learning Visual Representations from Video.
- Follow-up: V-JEPA 2 (2025). arXiv:2506.09985.
VL-JEPA (vision + language)
VL-JEPA replaces the standard autoregressive token-generation objective of vision-language models with a JEPA-style objective. Given a video and a textual query, it predicts the embedding of the answer rather than decoding text from left to right. It uses a lightweight decoder only when it needs explicit text output. This approach cuts decoding cost by roughly 2.85× while matching or exceeding autoregressive VLM baselines with half the trainable parameters.- Paper: Meta AI (2025). VL-JEPA: Joint Embedding Predictive Architecture for Vision-language, arXiv.
H-JEPA (hierarchical world models)
H-JEPA refers to the approach to long-horizon planning in LeCun’s 2022 position paper rather than to a single paper. The approach stacks JEPAs. Lower levels predict fine-grained, short-horizon representations such as seconds of video and millisecond-scale dynamics. Higher levels predict coarse-grained, long-horizon representations such as plans over minutes or hours and abstract subgoals. An agent can plan at the top of the stack without simulating every pixel of every frame. MC-JEPA’s pyramidal flow predictor is an early implementation of this approach.- Source: Yann LeCun (2022). A Path Towards Autonomous Machine Intelligence, Section 4 (“Hierarchical JEPA”) of the position paper.
- Related implementation: Bardes, Ponce, LeCun (2023). MC-JEPA, hierarchical flow prediction pyramid.
LeJEPA (theoretical analysis)
LeJEPA analyzes why JEPAs work and how to train them without the empirical methods used by earlier variants, including stop-gradient, an EMA teacher, schedulers, and large-batch tuning. It proves that the optimal embedding distribution for a JEPA is an isotropic Gaussian. It also introduces Sketched Isotropic Gaussian Regularization (SIGReg), a regularizer that drives training toward this distribution. The resulting objective has one hyperparameter and linear cost. Its implementation takes about 50 lines of code and matches or exceeds baselines that use the earlier heuristics across more than 10 datasets and more than 60 architectures.- Paper: Balestriero & LeCun (2025). LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics, arXiv.
- Code: rbalestr-lab/lejepa.

