Related Experiment Video
Updated: Jun 16, 2026

Automated Visual Cognitive Tasks for Recording Neural Activity Using a Floor Projection Maze
Published on: February 20, 2014
Online learning of shaping rewards in reinforcement learning
1Department of Computer Science, University of York, York YO105DD, UK. grzes@cs.york.ac.uk
Abstract:
Potential-based reward shaping has been shown to be a powerful method to improve the convergence rate of reinforcement learning agents. It is a flexible technique to incorporate background knowledge into temporal-difference learning in a principled way. However, the question remains of how to compute the potential function which is used to shape the reward that is given to the learning agent. In this paper, we show how, in the absence of knowledge to define the potential function manually, this function can be learned online in parallel with the actual reinforcement learning process. Two cases are considered. The first solution which is based on the multi-grid discretisation is designed for model-free reinforcement learning. In the second case, the approach for the prototypical model-based R-max algorithm is proposed. It learns the potential function using the free space assumption about the transitions in the environment. Two novel algorithms are presented and evaluated.
Related Concept Videos
Role of Shaping in Operant Conditioning
The steps involved in shaping begin with reinforcing any response that resembles the desired behavior. For example, parents might praise a child for picking up one toy. As...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Observational Learning
Reinforcement Schedules
Once a behavior is learned,...
Operant Conditioning
Reinforcement in operant conditioning can be positive or negative, both of which serve to increase the likelihood of a behavior. Positive...
Law of Effect
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle boxes...

