Related Experiment Video
Updated: Nov 19, 2025

07:52
Investigating Motor Skill Learning Processes with a Robotic Manipulandum
Published on: February 12, 2017
9.0K
Structure-Preserving Imitation Learning With Delayed Reward: An Evaluation Within the RoboCup Soccer 2D Simulation
Quang Dang Nguyen1, Mikhail Prokopenko1
1Centre for Complex Systems, Faculty of Engineering, University of Sydney, Sydney, NSW, Australia.
Frontiers in Robotics and AI
|January 27, 2021
Summary
We developed a neural network to enhance autonomous soccer robots in RoboCup. This deep Q-network approach improves team performance and can integrate expert-designed behaviors.
Area of Science:
- Robotics
- Artificial Intelligence
- Machine Learning
Background:
- Fully autonomous soccer teams in RoboCup require sophisticated control architectures.
- Existing behavioral modules in teams like Gliders2d are often evolved with human expert input.
Purpose of the Study:
- To design and evaluate a neural network architecture for improving autonomous soccer team performance.
- To assess the feasibility of replacing existing modules with a neural network approach.
- To investigate the impact of preserving expert-designed structures on neural network performance.
Main Methods:
- Utilized a deep Q-network for action determination and a deep neural network for parameter learning.
- Integrated the neural network into the Gliders2d RoboCup team.
- Introduced a delayed reward signal to aid the training process.
Main Results:
- The neural network-based architecture proved feasible for module replacement in a complex team.
- The addition of a delayed reward signal facilitated training.
- Performance was evaluated against a known benchmark.
Conclusions:
- Neural network architectures, specifically deep Q-networks, can effectively enhance autonomous robotic soccer performance.
- The integration of expert-designed behavioral structures can be maintained within neural network solutions.
- The proposed methods offer a viable path for improving AI agents in simulated environments.
Related Concept Videos
Observational Learning
625
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
625
Reinforcement Schedules
327
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
327
Avoidance Learning and Learned Helplessness
2.3K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
2.3K

