Skip to main content
JEPA (Joint-Embedding Predictive Architecture) is Yann LeCun’s proposed basis for world models. A JEPA predicts a representation of a future observation from a representation of a past observation rather than predicting raw pixels or tokens. The learned latent space lets the predictor ignore unpredictable surface details such as pixel noise, lighting changes, and paraphrase variation. It can then focus on the structural and semantic parts of the signal that support planning. This section presents the JEPA variants in the order in which they were proposed.

JEPA (theory, 2022)

LeCun’s framing paper argues that generative prediction in observation space is the wrong objective for learning a world model. Such a model spends capacity on irrelevant pixel-level detail. LeCun proposes that perception, prediction, planning, and short- and long-term memory should instead operate on learned representations. Each later JEPA implements part of this design.

I-JEPA (images)

I-JEPA is the first concrete JEPA. Given a single image, it masks target blocks. A predictor operates on encoded patches rather than pixels and predicts the representations of the masked blocks from the unmasked context. I-JEPA matches masked-autoencoder quality on ImageNet without a pixel-level reconstruction loss. This result supports latent-space prediction for static images.

MC-JEPA (motion + content)

MC-JEPA extends I-JEPA beyond a single still image. It jointly learns content (what is in the scene) and motion (how it flows) from video pairs. A shared encoder produces representations for semantic tasks such as classification and segmentation, as well as for optical-flow estimation. The results show that the JEPA objective can learn temporal structure while retaining static-image quality.

V-JEPA (video / dynamics)

V-JEPA applies latent-space prediction to full video. Its encoder processes many seconds of video, and its predictor reconstructs the representations of masked space-time regions. The resulting features transfer to downstream video-understanding tasks without training with a generative pixel loss. V-JEPA 2 (2025) scales the model to billions of parameters and serves as the visual backbone for several later VLA and VL-JEPA systems.

VL-JEPA (vision + language)

VL-JEPA replaces the standard autoregressive token-generation objective of vision-language models with a JEPA-style objective. Given a video and a textual query, it predicts the embedding of the answer rather than decoding text from left to right. It uses a lightweight decoder only when it needs explicit text output. This approach cuts decoding cost by roughly 2.85× while matching or exceeding autoregressive VLM baselines with half the trainable parameters.

H-JEPA (hierarchical world models)

H-JEPA refers to the approach to long-horizon planning in LeCun’s 2022 position paper rather than to a single paper. The approach stacks JEPAs. Lower levels predict fine-grained, short-horizon representations such as seconds of video and millisecond-scale dynamics. Higher levels predict coarse-grained, long-horizon representations such as plans over minutes or hours and abstract subgoals. An agent can plan at the top of the stack without simulating every pixel of every frame. MC-JEPA’s pyramidal flow predictor is an early implementation of this approach.

LeJEPA (theoretical analysis)

LeJEPA analyzes why JEPAs work and how to train them without the empirical methods used by earlier variants, including stop-gradient, an EMA teacher, schedulers, and large-batch tuning. It proves that the optimal embedding distribution for a JEPA is an isotropic Gaussian. It also introduces Sketched Isotropic Gaussian Regularization (SIGReg), a regularizer that drives training toward this distribution. The resulting objective has one hyperparameter and linear cost. Its implementation takes about 50 lines of code and matches or exceeds baselines that use the earlier heuristics across more than 10 datasets and more than 60 architectures.

JEPA in world models

The World Models section considers what an internal simulator should predict so that an agent can plan at human-relevant time scales. The JEPA family predicts representations rather than pixels or tokens. The papers above apply this idea to static images (I-JEPA), multiple modalities (VL-JEPA), hierarchical models (H-JEPA), and theoretical analysis (LeJEPA).