Artificial intelligence
Offline reinforcement learning
Definition
Offline reinforcement learning learns a reward-optimizing policy from previously collected experience without gathering new environment interactions during that learning stage. The data may come from earlier policies, demonstrations, or other collection procedures.
Also known as: Offline RL, Batch reinforcement learning
Updated
Reuse an existing experience dataset
The offline reinforcement-learning tutorial defines the setting around learning without additional online data collection. Recorded transitions describe what the agent observed, which action it took, what happened next, and the associated reward.
A robotics lab might reuse earlier manipulation trials rather than repeatedly running a new learner on hardware. That changes the learning problem because the dataset cannot automatically supply examples for every action the learner now wants to try.
Unseen actions can receive misleading values
A policy may favor actions for which the dataset contains little support. A learned value function or dynamics model can make optimistic predictions about those actions, and there is no new interaction during offline training to correct the error.
The tutorial discusses distribution shift, conservative value estimation, and constraints that keep learned behavior closer to the data as responses to this problem.
Offline does not mean behavior cloning
Behavior cloning learns to reproduce recorded actions. Offline reinforcement learning tries to optimize return from the available experience, potentially choosing different behavior.
Dataset coverage and reward quality limit what can be inferred. Offline training also does not remove the need to evaluate the resulting policy. A later online fine-tuning stage is possible, but it is a separate stage from the offline learning described here.
Sources
Related terms
Reinforcement learning
Reinforcement learning trains an agent to choose actions that maximize expected cumulative reward through experience with an environment. In robotics, the learned policy can select movements or higher-level behaviors from observations.
Behavior cloning
Behavior cloning is an imitation-learning method that trains a policy to predict a demonstrator’s actions from recorded observations or states. It treats action prediction as a supervised-learning problem.
World model
A world model is an internal predictive model of an environment and how it changes. In robot learning, it can predict future states or observations under possible actions to support planning or policy training.