Skip to main content
The original World Models paper by Ha & Schmidhuber (2018) showed that an agent can learn to act inside a learned simulator of its environment. Its architecture provides a useful starting point for studying modern world models. World model agent driving in CarRacing, the agent navigates based on its learned internal model of the environment

The core architecture

The agent has three components: V, Vision model (VAE). A variational autoencoder compresses each high-dimensional observation (e.g., a 64x64 RGB frame from CarRacing) into a latent vector z. It reduces a 12,288-dimensional pixel input to about 32 latent dimensions while preserving the information needed for control. M, Memory model (MDN-RNN). A recurrent neural network with a mixture density output predicts the next latent state from the current latent state and action. This model captures the environment’s dynamics in latent space. Its mixture density network (MDN) output represents uncertainty by predicting a distribution over possible next states. C, Controller. A small linear controller maps the current latent state z and the RNN hidden state h to an action. Because V and M have already compressed the observation and learned the dynamics, the controller can be very simple, often just a single linear layer optimized with evolutionary strategies (CMA-ES).

Separation of components

Each component has a separate role:
  • V reduces the input dimension, so the controller never sees raw pixels.
  • M predicts temporal changes, so the controller does not need to learn the dynamics.
  • C selects actions in a low-dimensional, temporally structured space.
The controller can therefore be trained inside the world model without interacting with the environment. Dream trajectories are generated by rolling out M from a starting state. C is then optimized against those trajectories. In this sense, the agent learns to drive by dreaming about driving.

The CarRacing experiment

The original paper demonstrates the approach on CarRacing. The behavioral cloning tutorial uses the same environment, so you can compare BC, PPO, and the world model on one task. Training proceeds in five stages:
  1. Collect data. Run a random or partially trained policy in CarRacing to collect 10,000 frames of (observation, action, next_observation) tuples.
  2. Train V. Train the VAE on the collected frames to learn the latent representation.
  3. Train M. Train the MDN-RNN on sequences of (z, action, z_next) to learn the dynamics model.
  4. Train C. Use CMA-ES to evolve the controller parameters by evaluating candidate controllers inside M’s dream rollouts. This stage does not require interaction with the environment.
  5. Deploy. Run V + M + C in the real environment: observe → encode → predict → act.

Further reading