
The core architecture
The agent has three components: V, Vision model (VAE). A variational autoencoder compresses each high-dimensional observation (e.g., a 64x64 RGB frame from CarRacing) into a latent vector z. It reduces a 12,288-dimensional pixel input to about 32 latent dimensions while preserving the information needed for control. M, Memory model (MDN-RNN). A recurrent neural network with a mixture density output predicts the next latent state from the current latent state and action. This model captures the environment’s dynamics in latent space. Its mixture density network (MDN) output represents uncertainty by predicting a distribution over possible next states. C, Controller. A small linear controller maps the current latent state z and the RNN hidden state h to an action. Because V and M have already compressed the observation and learned the dynamics, the controller can be very simple, often just a single linear layer optimized with evolutionary strategies (CMA-ES).Separation of components
Each component has a separate role:- V reduces the input dimension, so the controller never sees raw pixels.
- M predicts temporal changes, so the controller does not need to learn the dynamics.
- C selects actions in a low-dimensional, temporally structured space.
The CarRacing experiment
The original paper demonstrates the approach on CarRacing. The behavioral cloning tutorial uses the same environment, so you can compare BC, PPO, and the world model on one task. Training proceeds in five stages:- Collect data. Run a random or partially trained policy in CarRacing to collect 10,000 frames of (observation, action, next_observation) tuples.
- Train V. Train the VAE on the collected frames to learn the latent representation.
- Train M. Train the MDN-RNN on sequences of (z, action, z_next) to learn the dynamics model.
- Train C. Use CMA-ES to evolve the controller parameters by evaluating candidate controllers inside M’s dream rollouts. This stage does not require interaction with the environment.
- Deploy. Run V + M + C in the real environment: observe → encode → predict → act.
Further reading
- Ha & Schmidhuber (2018). World Models, the original paper with interactive visualizations
- Ha & Schmidhuber (2018). Recurrent World Models Facilitate Policy Evolution, the NeurIPS version
- An implementation for reproducing the CarRacing experiment

