Artificial intelligence
World model
Definition
A world model is an internal predictive model of an environment and how it changes. In robot learning, it can predict future states or observations under possible actions to support planning or policy training.
Also known as: World models
Updated
Predict what may happen next
Ha and Schmidhuber's World Models learns a compressed representation of observations and a recurrent model of how that representation evolves. A controller can use features from the model, and the authors also train an agent inside a generated environment.
For a robot, a corresponding use would be predicting how an object might move after a push. The prediction could concern a compact state representation instead of a photorealistic future image.
Prediction and control remain different roles
A world model predicts outcomes; a policy chooses actions. One system can contain both, but a predictive video model does not automatically provide a robot controller.
The cited work evaluates simulated reinforcement-learning environments. Its generated-environment training demonstrates the method in that setting, rather than establishing that any visually plausible model accurately predicts physical robot contact.
Controllers can exploit model errors
The authors describe a controller finding behavior that exploits mistakes in its learned environment. Those opportunities did not behave the same way in the original environment.
This makes model fidelity relevant to the task and to the actions considered. A model that reconstructs familiar observations well can still make poor predictions for unfamiliar actions. System identification and world-model learning overlap where both estimate dynamics from data, although world-model architectures need not use explicit mechanical equations.
Sources
Related terms
Model predictive control
Model predictive control repeatedly optimizes future actions using a system model, applies the next part of the solution, and replans from updated state information. It can account for objectives and constraints over a finite prediction horizon.
Reinforcement learning
Reinforcement learning trains an agent to choose actions that maximize expected cumulative reward through experience with an environment. In robotics, the learned policy can select movements or higher-level behaviors from observations.
System identification
System identification estimates a model of a physical system from measured inputs and outputs. In robotics, it can recover parameters such as inertia and friction or learn a more general model of how actions change the system state.