World models
A world model is a learned function of the form- Plan by searching over hypothetical action sequences.
- Train a policy on imagined rollouts without collecting real-world samples.
- Estimate uncertainty by sampling several futures from the model.
A family of architectures
World models use several types of generative backbone. Each type follows the same basic method. The agent learns an internal simulator and trains a controller inside it. Newer architectures have increased the fidelity and scope of the simulation.
The original VAE-based architecture produces blurry reconstructions and uses a fixed latent dimension. Its simple structure makes it a useful introduction to larger variants. Diffusion and autoregressive models produce images with higher fidelity, but they use many more parameters and generate rollouts more slowly. The JEPA family takes a different approach. It predicts representations of future observations instead of reconstructing pixels or tokens. This lets the model focus on semantically meaningful structure.
Comparison with behavioral cloning
World models and behavioral cloning learn control policies in different ways. They also have different failure modes.World models in Physical AI
World models relate to several topics in this course. Sim-to-real. A world model trained on real data is a simulator calibrated from real observations. This removes the hand-authored simulation gap discussed in the sim-to-real transfer section. You can learn a world model from a few minutes of real video and train in it instead of building a Gazebo world by hand. Behavioral cloning and world models. BC uses expert demonstrations as its only data source. A world model learns the environment dynamics and can extrapolate beyond those demonstrations by predicting outcomes in unseen states. This can reduce BC’s distribution-shift problem because the agent trains on a wider distribution of states than a fixed demonstration dataset provides. VLA models. Most current VLA architectures, such as OpenVLA and RT-2, use behavioral cloning without an explicit world model. Combining these architectures with learned dynamics models is an active research area. Such combinations may reduce the limitations of behavioral cloning described above.Model families
The Seminal Model
Ha and Schmidhuber’s original V/M/C architecture uses a VAE for vision, an MDN-RNN for memory, and a small linear controller trained inside the model.
The JEPA family
LeCun’s roadmap for world models predicts representations instead of pixels. It covers the 2022 position paper, I-JEPA, MC-JEPA, V-JEPA, VL-JEPA, H-JEPA, and LeJEPA.
Further reading
- Ha & Schmidhuber (2018). World Models. This paper introduced the original architecture.
- Hafner et al. (2020). Dream to Control: Learning Behaviors by Latent Imagination. Dreamer uses actor-critic training in latent space.
- Hafner et al. (2023). Mastering Diverse Domains through World Models. DreamerV3 uses the same architecture across Atari, DMControl, and Minecraft.
- Alonso et al. (2024). Diffusion for World Modeling: Visual Details Matter in Atari. This paper introduces DIAMOND.
- Hu et al. (2023). GAIA-1: A Generative World Model for Autonomous Driving

