Skip to main content
Humans form internal models of the world and use them to update their expectations. When you catch a ball, you do not react separately to each frame of visual input. You predict the ball’s trajectory and plan your arm’s motion from that prediction. World models apply this idea to machine learning. Instead of learning a policy directly from interactions with the environment (model-free RL), an agent first learns a compressed predictive model of the environment. It then trains a controller inside that model.

World models

A world model is a learned function of the form
The agent can roll this function forward in time without interacting with the real environment. You can use it to:
  • Plan by searching over hypothetical action sequences.
  • Train a policy on imagined rollouts without collecting real-world samples.
  • Estimate uncertainty by sampling several futures from the model.
This is a deep-learning form of model-based reinforcement learning. Classical model-based RL uses a hand-coded transition function, as in Dyna and MPC. Modern world models learn the transition function directly from pixels.

A family of architectures

World models use several types of generative backbone. Each type follows the same basic method. The agent learns an internal simulator and trains a controller inside it. Newer architectures have increased the fidelity and scope of the simulation. The original VAE-based architecture produces blurry reconstructions and uses a fixed latent dimension. Its simple structure makes it a useful introduction to larger variants. Diffusion and autoregressive models produce images with higher fidelity, but they use many more parameters and generate rollouts more slowly. The JEPA family takes a different approach. It predicts representations of future observations instead of reconstructing pixels or tokens. This lets the model focus on semantically meaningful structure.

Comparison with behavioral cloning

World models and behavioral cloning learn control policies in different ways. They also have different failure modes.

World models in Physical AI

World models relate to several topics in this course. Sim-to-real. A world model trained on real data is a simulator calibrated from real observations. This removes the hand-authored simulation gap discussed in the sim-to-real transfer section. You can learn a world model from a few minutes of real video and train in it instead of building a Gazebo world by hand. Behavioral cloning and world models. BC uses expert demonstrations as its only data source. A world model learns the environment dynamics and can extrapolate beyond those demonstrations by predicting outcomes in unseen states. This can reduce BC’s distribution-shift problem because the agent trains on a wider distribution of states than a fixed demonstration dataset provides. VLA models. Most current VLA architectures, such as OpenVLA and RT-2, use behavioral cloning without an explicit world model. Combining these architectures with learned dynamics models is an active research area. Such combinations may reduce the limitations of behavioral cloning described above.

Model families

The Seminal Model

Ha and Schmidhuber’s original V/M/C architecture uses a VAE for vision, an MDN-RNN for memory, and a small linear controller trained inside the model.

The JEPA family

LeCun’s roadmap for world models predicts representations instead of pixels. It covers the 2022 position paper, I-JEPA, MC-JEPA, V-JEPA, VL-JEPA, H-JEPA, and LeJEPA.

Further reading