Artificial intelligence
Reward shaping
Definition
Reward shaping adds supplementary rewards to guide reinforcement learning toward useful behavior. Poorly chosen shaping can change which policy is optimal, so an easier training signal is not automatically equivalent to the original task objective.
Updated
Provide feedback before task completion
A robot may receive its main reward only after reaching a goal. Shaping can provide intermediate feedback, such as a signal related to progress toward that goal. Ng, Harada, and Russell analyze when adding such feedback preserves the original optimal policy.
The distinction matters because a learner optimizes the reward it actually receives, including any extra terms.
Potential-based shaping has a precise form
For a discounted problem, potential-based shaping adds gamma * Phi(next_state) - Phi(state), where Phi assigns a potential to each state and gamma is the discount factor. The paper establishes policy invariance under its stated Markov decision process assumptions and boundary conditions.
This form rewards a change in potential rather than handing out an unrelated bonus whenever a convenient event occurs.
Extra rewards can create unintended loops
The paper describes examples in which rewarding progress or repeated ball contact encouraged behavior that failed the intended objective. For a robot moving a part, repeatedly collecting a proximity bonus could likewise become preferable to completing placement if the reward is poorly specified.
Shaping is a tool within reinforcement learning, not a substitute for defining the task. Report the full reward and verify actual task completion rather than relying only on the shaped return.
Sources
Related terms
Reinforcement learning
Reinforcement learning trains an agent to choose actions that maximize expected cumulative reward through experience with an environment. In robotics, the learned policy can select movements or higher-level behaviors from observations.
Hierarchical reinforcement learning
Hierarchical reinforcement learning organizes learned decision-making into levels, often with a higher-level policy selecting goals or skills and lower-level policies producing actions. The levels can operate over different time scales.
Residual reinforcement learning
Residual reinforcement learning learns a corrective control signal that is combined with a baseline controller. The baseline handles part of the task while the learned residual adjusts behavior that is difficult to model or tune directly.