Artificial intelligence

Offline reinforcement learning

Definition

Offline reinforcement learning learns a reward-optimizing policy from previously collected experience without gathering new environment interactions during that learning stage. The data may come from earlier policies, demonstrations, or other collection procedures.

Also known as: Offline RL, Batch reinforcement learning

Updated

Reuse an existing experience dataset

The offline reinforcement-learning tutorial defines the setting around learning without additional online data collection. Recorded transitions describe what the agent observed, which action it took, what happened next, and the associated reward.

A robotics lab might reuse earlier manipulation trials rather than repeatedly running a new learner on hardware. That changes the learning problem because the dataset cannot automatically supply examples for every action the learner now wants to try.

Unseen actions can receive misleading values

A policy may favor actions for which the dataset contains little support. A learned value function or dynamics model can make optimistic predictions about those actions, and there is no new interaction during offline training to correct the error.

The tutorial discusses distribution shift, conservative value estimation, and constraints that keep learned behavior closer to the data as responses to this problem.

Offline does not mean behavior cloning

Behavior cloning learns to reproduce recorded actions. Offline reinforcement learning tries to optimize return from the available experience, potentially choosing different behavior.

Dataset coverage and reward quality limit what can be inferred. Offline training also does not remove the need to evaluate the resulting policy. A later online fine-tuning stage is possible, but it is a separate stage from the offline learning described here.

Sources