Related Experiment Video
Updated: Oct 22, 2025

06:57
Pavlovian Conditioned Approach Training in Rats
Published on: February 4, 2016
11.1K
A Multi-Dimensional Goal Aircraft Guidance Approach Based on Reinforcement Learning with a Reward Shaping Algorithm
Wenqiang Zu1, Hongyu Yang1, Renyu Liu1
1College of Computer Science, Sichuan University, Chengdu 610065, China.
Sensors (Basel, Switzerland)
|August 28, 2021
Summary
This study introduces a multi-layer reinforcement learning (RL) approach for aircraft guidance to 4D waypoints. The method enhances convergence and trajectory performance in air traffic control simulations.
Area of Science:
- Aerospace Engineering
- Artificial Intelligence
- Control Systems
Background:
- Aircraft guidance to 4D waypoints presents a complex multi-dimensional control challenge.
- Current air traffic control (ATC) simulators face limitations in managing direct aircraft approach patterns.
Purpose of the Study:
- To propose and evaluate a multi-layer reinforcement learning (RL) approach for solving the multi-dimensional goal aircraft guidance problem.
- To enhance the performance and efficiency of aircraft guidance systems within ATC simulators.
Main Methods:
- A multi-layer RL method is developed to simplify neural network architecture and reduce state dimensions.
- A shaped reward function incorporating potential functions and the Dubins path method is applied.
- The approach assists autopilot systems in guiding aircraft to specified 4D waypoints (latitude, longitude, altitude, heading, arrival time).
Main Results:
- Simulation results demonstrate significant improvements in convergence efficiency compared to existing methods.
- The proposed approach shows enhanced trajectory performance for aircraft guidance.
- Experimental validation confirms the effectiveness of the multi-layer RL strategy.
Conclusions:
- The multi-layer RL approach effectively addresses the complexities of 4D waypoint guidance.
- The method offers potential applications in team aircraft guidance, enabling direct goal approaches.
- This research overcomes limitations in current ATC simulators by facilitating more efficient aircraft routing.
Related Concept Videos
Role of Shaping in Operant Conditioning
591
Shaping is a technique used in operant conditioning to train complex behaviors by rewarding successive approximations toward the target behavior. This method is necessary because organisms are unlikely to perform complex behaviors spontaneously. Instead, shaping breaks down the desired behavior into small, manageable steps.
The steps involved in shaping begin with reinforcing any response that resembles the desired behavior. For example, parents might praise a child for picking up one toy. As...
The steps involved in shaping begin with reinforcing any response that resembles the desired behavior. For example, parents might praise a child for picking up one toy. As...
591
Reinforcement
483
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
483
Observational Learning
405
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
405
Avoidance Learning and Learned Helplessness
2.0K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
2.0K
Reinforcement Schedules
275
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
275
Multi-input and Multi-variable systems
206
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
In the absence...
206

