Artificial intelligence
Residual reinforcement learning
Definition
Residual reinforcement learning learns a corrective control signal that is combined with a baseline controller. The baseline handles part of the task while the learned residual adjusts behavior that is difficult to model or tune directly.
Also known as: Residual RL
Updated
Learn a correction around an existing controller
Residual Reinforcement Learning for Robot Control decomposes control into a conventional feedback-control component and a learned residual. In the paper's formulation, the final command is the superposition of the two signals.
A baseline controller could guide an arm toward a target while the residual adjusts the command during contact. The learned component does not have to rediscover every aspect of the baseline behavior.
Contact is a motivating application
The authors identify contacts and friction as effects that can be difficult to capture with simple physical models. They demonstrate the approach on a real block-assembly task involving contacts and unstable objects.
This is a specific use of reinforcement learning with an existing controller. The word residual refers to the correction added to the control signal, not to a particular residual neural-network architecture.
A baseline does not constrain every correction
Adding a learned signal changes the command the robot executes. The baseline's behavior alone therefore does not establish the behavior or stability of the combined system.
The baseline, residual action space, allowed correction magnitude, and task evaluation all matter when assessing an implementation. Results on block assembly support that demonstrated application; they do not imply that adding an arbitrary learned correction will improve every controller or preserve its guarantees.
Sources
Related terms
Reinforcement learning
Reinforcement learning trains an agent to choose actions that maximize expected cumulative reward through experience with an environment. In robotics, the learned policy can select movements or higher-level behaviors from observations.
Force control
Force control regulates the force or wrench a robot applies to its environment. It may use a robot model, measured interaction forces, or both to produce joint commands that achieve a desired contact load.
Impedance control
Impedance control shapes the dynamic relationship between a robot’s motion and the forces it exchanges with its environment. A common goal is for the robot to respond like a chosen mass, spring, and damper at a joint or end effector.