π World Model
Descriptionβ
< What is it? >β
A world model is a learned approximation of an environment. Given the current observation or state and an action, it predicts what may happen nextβcommonly a next latent state and a reward. In model-based reinforcement learning, an agent can plan or improve its policy using predicted, or imagined, trajectories instead of relying only on real interactions.
- Example: before a robot drives forward, its world model can predict that the action will bring it closer to a wall and likely incur a collision penalty. The robot can consider that outcome before taking the physical action.
Key pointsβ
< What does it learn? >β
Many world models first compress a raw observation , such as an image, into a compact latent state . A dynamics model then predicts the next latent state and reward after action :
Here, is an encoder, is the learned dynamics-and-reward model, and hats denote predictions. Some world models also predict whether an episode will end or reconstruct an observation from the latent state.
< How does an agent use it? >β
The agent can roll the model forward repeatedly to compare possible action sequences:
A planner can choose the sequence with the best predicted return. Alternatively, an actor and value function can be trained on imagined rollouts, then the resulting policy is used in the real environment.
< Why use a latent state? >β
Predicting every pixel of a future image can be expensive and can force the model to preserve details that do not matter for decisions. A latent state aims to retain information useful for predicting rewards and consequences. A world model therefore does not need to be a visually perfect simulator.
MuZero is an important example: it learns an internal model that predicts planning-relevant reward, value, and policy quantities rather than being given the environment's rules.
< Benefits and limitations >β
- Potential benefit: imagined rollouts can make learning more sample-efficient and allow an agent to look ahead before acting.
- Main limitation: small prediction errors can compound over a long imagined rollout. A policy may exploit inaccuracies in its model instead of behaving well in the real environment.
- Common response: use short rollouts, replan from new real observations, model uncertainty, and continue collecting real data.
Comparisonβ
< World-model RL vs model-free RL >β
| World-model reinforcement learning | Model-free reinforcement learning | |
|---|---|---|
| What is learned | A predictive model of consequences, often plus a policy or value function | A policy, value function, or both directly from experience |
| Action selection | Can plan or learn from imagined futures | Uses the learned policy or value estimates without an explicit learned dynamics model |
| Potential strength | Better sample efficiency and look-ahead when the model is accurate | Simpler objective; avoids errors from a learned simulator |
| Main risk | Model errors compound or are exploited during planning | May require many real environment interactions |
| Examples | Dreamer, MuZero | DQN, standard policy-gradient methods |
Video Tutorialβ
- Googleβs Frontier World Model (Genie 3), Explained in 2 Mins
- AI World Models Explained
- World Models
Related ideasβ
- Deep Q Network (DQN) is a model-free value-based method.
- Policy Gradient introduces policy optimization; some world-model methods optimize policies on imagined trajectories.
- Embeddings explains the compact vector representations that a latent state resembles.
Referenceβ
- World Models β David Ha and JΓΌrgen Schmidhuber
- Mastering Diverse Domains through World Models β DreamerV3
- Mastering Atari, Go, chess and shogi by planning with a learned model β MuZero