Skip to main content

πŸ“ World Model

Description​

< What is it? >​

A world model is a learned approximation of an environment. Given the current observation or state and an action, it predicts what may happen nextβ€”commonly a next latent state and a reward. In model-based reinforcement learning, an agent can plan or improve its policy using predicted, or imagined, trajectories instead of relying only on real interactions.

  • Example: before a robot drives forward, its world model can predict that the action will bring it closer to a wall and likely incur a collision penalty. The robot can consider that outcome before taking the physical action.

Key points​

< What does it learn? >​

Many world models first compress a raw observation oto_t, such as an image, into a compact latent state ztz_t. A dynamics model then predicts the next latent state and reward after action ata_t:

zt=eΟ•(ot),(z^t+1,r^t)=fΞΈ(zt,at)z_t=e_\phi(o_t), \qquad (\hat z_{t+1},\hat r_t)=f_\theta(z_t,a_t)

Here, eΟ•e_\phi is an encoder, fΞΈf_\theta is the learned dynamics-and-reward model, and hats denote predictions. Some world models also predict whether an episode will end or reconstruct an observation from the latent state.

< How does an agent use it? >​

The agent can roll the model forward repeatedly to compare possible action sequences:

zt→atz^t+1→at+1z^t+2z_t\xrightarrow{a_t}\hat z_{t+1}\xrightarrow{a_{t+1}}\hat z_{t+2}

A planner can choose the sequence with the best predicted return. Alternatively, an actor and value function can be trained on imagined rollouts, then the resulting policy is used in the real environment.

< Why use a latent state? >​

Predicting every pixel of a future image can be expensive and can force the model to preserve details that do not matter for decisions. A latent state aims to retain information useful for predicting rewards and consequences. A world model therefore does not need to be a visually perfect simulator.

MuZero is an important example: it learns an internal model that predicts planning-relevant reward, value, and policy quantities rather than being given the environment's rules.

< Benefits and limitations >​

  • Potential benefit: imagined rollouts can make learning more sample-efficient and allow an agent to look ahead before acting.
  • Main limitation: small prediction errors can compound over a long imagined rollout. A policy may exploit inaccuracies in its model instead of behaving well in the real environment.
  • Common response: use short rollouts, replan from new real observations, model uncertainty, and continue collecting real data.

Comparison​

< World-model RL vs model-free RL >​

World-model reinforcement learningModel-free reinforcement learning
What is learnedA predictive model of consequences, often plus a policy or value functionA policy, value function, or both directly from experience
Action selectionCan plan or learn from imagined futuresUses the learned policy or value estimates without an explicit learned dynamics model
Potential strengthBetter sample efficiency and look-ahead when the model is accurateSimpler objective; avoids errors from a learned simulator
Main riskModel errors compound or are exploited during planningMay require many real environment interactions
ExamplesDreamer, MuZeroDQN, standard policy-gradient methods

Video Tutorial​

  • Deep Q Network (DQN) is a model-free value-based method.
  • Policy Gradient introduces policy optimization; some world-model methods optimize policies on imagined trajectories.
  • Embeddings explains the compact vector representations that a latent state resembles.

Reference​