Related Experiment Video
Updated: May 7, 2026

05:41
A Step-by-Step Implementation of DeepBehavior, Deep Learning Toolbox for Automated Behavior Analysis
Published on: February 6, 2020
Batch Mode Reinforcement Learning based on the Synthesis of Artificial Trajectories
Raphael Fonteneau1, Susan A Murphy, Louis Wehenkel
1University of Liége, Belgium.
Summary
This study introduces artificial trajectories to improve batch mode reinforcement learning in continuous state spaces. This novel approach offers new methods for designing and analyzing reinforcement learning algorithms.
Area of Science:
- Artificial intelligence
- Machine learning
- Control theory
Background:
- Batch mode reinforcement learning (RL) aims to derive optimal policies from pre-collected trajectory data.
- Continuous state-space problems in RL typically require function approximators for control or value functions.
- Existing methods face challenges in efficiently learning from limited or discrete trajectory samples.
Purpose of the Study:
- To propose an alternative to function approximators in batch mode reinforcement learning for continuous state spaces.
- To introduce the concept of synthesizing "artificial trajectories" from existing data.
- To demonstrate the potential of this new approach for designing and analyzing RL algorithms.
Main Methods:
- Focus on the batch mode reinforcement learning setting with continuous state spaces.
- Synthesize "artificial trajectories" using the provided sample of trajectories.
- Develop and analyze new algorithms based on the artificial trajectory concept.
Main Results:
- Demonstrated that artificial trajectories can be effectively synthesized from sample data.
- Showcased that this method provides a viable alternative to traditional function approximation techniques.
- Opened new research avenues for designing and analyzing batch mode reinforcement learning algorithms.
Conclusions:
- The synthesis of artificial trajectories presents a promising direction for advancing batch mode reinforcement learning.
- This technique offers a novel perspective for overcoming limitations associated with function approximators in continuous state spaces.
- Further research into artificial trajectory generation can lead to more robust and efficient RL algorithms.
Related Concept Videos
Observational Learning
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning because...
Synthetic Biology
Synthetic biology is an interdisciplinary science that involves using principles from disciplines such as engineering, molecular biology, cell biology, and systems biology. It involves remodeling existing organisms from nature or constructing completely new synthetic organisms for applications such as protein or enzyme production, bioremediation, value-added macromolecule production, and the addition of desirable traits to crops, to name a few.
Golden rice
Golden rice is a genetically modified...
Golden rice
Golden rice is a genetically modified...
Orthogonal Trajectories
Orthogonal trajectories describe the geometric relationship between two families of curves that intersect each other at right angles. One illustrative case involves a family of parabolas that open sideways along the x-axis. These curves share a common shape but differ by a scaling parameter, resulting in a set of curves that all pass through the origin and widen at different rates.Determining Orthogonal TrajectoriesTo identify the orthogonal trajectories for these parabolas, the first step...
Role of Shaping in Operant Conditioning
Shaping is a technique used in operant conditioning to train complex behaviors by rewarding successive approximations toward the target behavior. This method is necessary because organisms are unlikely to perform complex behaviors spontaneously. Instead, shaping breaks down the desired behavior into small, manageable steps.
The steps involved in shaping begin with reinforcing any response that resembles the desired behavior. For example, parents might praise a child for picking up one toy. As...
The steps involved in shaping begin with reinforcing any response that resembles the desired behavior. For example, parents might praise a child for picking up one toy. As...
Reinforcement Schedules
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
Steps in the Modeling Process
Albert Bandura's theory of observational learning identifies four critical processes: attention, retention, motor reproduction, and reinforcement or motivation.
Attention is the first necessary component for observational learning. It involves focusing on what the model is doing and saying. For example, if you decide to take a drawing class to enhance your skills, you need to pay close attention to the instructor's words and hand movements. The characteristics of the model significantly...
Attention is the first necessary component for observational learning. It involves focusing on what the model is doing and saying. For example, if you decide to take a drawing class to enhance your skills, you need to pay close attention to the instructor's words and hand movements. The characteristics of the model significantly...