Artificial intelligence

Residual reinforcement learning

Definition

Residual reinforcement learning learns a corrective control signal that is combined with a baseline controller. The baseline handles part of the task while the learned residual adjusts behavior that is difficult to model or tune directly.

Also known as: Residual RL

Updated

Learn a correction around an existing controller

Residual Reinforcement Learning for Robot Control decomposes control into a conventional feedback-control component and a learned residual. In the paper's formulation, the final command is the superposition of the two signals.

A baseline controller could guide an arm toward a target while the residual adjusts the command during contact. The learned component does not have to rediscover every aspect of the baseline behavior.

Contact is a motivating application

The authors identify contacts and friction as effects that can be difficult to capture with simple physical models. They demonstrate the approach on a real block-assembly task involving contacts and unstable objects.

This is a specific use of reinforcement learning with an existing controller. The word residual refers to the correction added to the control signal, not to a particular residual neural-network architecture.

A baseline does not constrain every correction

Adding a learned signal changes the command the robot executes. The baseline's behavior alone therefore does not establish the behavior or stability of the combined system.

The baseline, residual action space, allowed correction magnitude, and task evaluation all matter when assessing an implementation. Results on block assembly support that demonstrated application; they do not imply that adding an arbitrary learned correction will improve every controller or preserve its guarantees.

Sources