Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Reinforcement Schedules01:24

Reinforcement Schedules

147
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
147
PD Controller: Design01:26

PD Controller: Design

229
In automotive engineering, car suspension systems often employ Proportional Derivative (PD) controllers to enhance performance. PD controllers are utilized to adjust the damping force in response to road conditions. A controller, acting as an amplifier with a constant gain, demonstrates proportional control, with output directly mirroring input.
Designing a continuous-data controller requires selecting and linking components like adders and integrators, which are fundamental in Proportional,...
229
Timing and Consequences on Behavior01:08

Timing and Consequences on Behavior

94
In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective. 
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
94
Time-Domain Interpretation of PD Control01:07

Time-Domain Interpretation of PD Control

111
Proportional-Derivative (PD) control is a widely used control method in various engineering systems to enhance stability and performance. In a system with only proportional control, common issues include high maximum overshoot and oscillation, observed in both the error signal and its rate of change. This behavior can be divided into three distinct phases: initial overshoot, subsequent undershoot, and gradual stabilization.
Consider the example of control of motor torque. Initially, a positive...
111
Feedback control systems01:26

Feedback control systems

313
Feedback control systems are categorized in various ways based on their design, analysis, and signal types.
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
313

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

[Correlations between DNA mismatch repair (MMR) and prognosis and prediction of treatment efficacy in stage II/II colon cancer].

Zhonghua zhong liu za zhi [Chinese journal of oncology]·2015
Same author

[Clinical effect of hemoperfusion combined with hemodialysis in treatment of severe organophosphate pesticide poisoning].

Zhonghua lao dong wei sheng zhi ye bing za zhi = Zhonghua laodong weisheng zhiyebing zazhi = Chinese journal of industrial hygiene and occupational diseases·2015
Same author

Effect of orexin A on apoptosis in BGC-823 gastric cancer cells via OX1R through the AKT signaling pathway.

Molecular medicine reports·2015
Same author

Time-resolved dynamic dilution introduction for ion mobility spectrometry and its application in end-tidal propofol monitoring.

Journal of breath research·2015
Same author

Curative effect assessment of bandage contact lens in neurogenic keratitis.

International journal of ophthalmology·2014
Same author

Orexin A upregulates the protein expression of OX1R and enhances the proliferation of SGC-7901 gastric cancer cells through the ERK signaling pathway.

International journal of molecular medicine·2014

Related Experiment Video

Updated: Jul 2, 2025

WheelCon: A Wheel Control-Based Gaming Platform for Studying Human Sensorimotor Control
08:18

WheelCon: A Wheel Control-Based Gaming Platform for Studying Human Sensorimotor Control

Published on: August 15, 2020

5.0K

Hierarchical Reinforcement Learning for UAV-PE Game With Alternative Delay Update Method.

Xiao Ma, Yuan Yuan, Lei Guo

    IEEE Transactions on Neural Networks and Learning Systems
    |February 21, 2024
    PubMed
    Summary

    This study introduces a new hierarchical reinforcement learning (HRL) algorithm with an alternative delay update (ADU) method for unmanned aerial vehicle pursuit-evasion (UAV-PE) games. The approach efficiently finds approximate Nash equilibrium solutions by coupling kinematics and dynamics.

    More Related Videos

    Recording Single Neurons' Action Potentials from Freely Moving Pigeons Across Three Stages of Learning
    11:20

    Recording Single Neurons' Action Potentials from Freely Moving Pigeons Across Three Stages of Learning

    Published on: June 2, 2014

    12.0K
    A Method for Investigating Change Blindness in Pigeons Columba Livia
    06:14

    A Method for Investigating Change Blindness in Pigeons Columba Livia

    Published on: September 7, 2018

    6.4K

    Related Experiment Videos

    Last Updated: Jul 2, 2025

    WheelCon: A Wheel Control-Based Gaming Platform for Studying Human Sensorimotor Control
    08:18

    WheelCon: A Wheel Control-Based Gaming Platform for Studying Human Sensorimotor Control

    Published on: August 15, 2020

    5.0K
    Recording Single Neurons' Action Potentials from Freely Moving Pigeons Across Three Stages of Learning
    11:20

    Recording Single Neurons' Action Potentials from Freely Moving Pigeons Across Three Stages of Learning

    Published on: June 2, 2014

    12.0K
    A Method for Investigating Change Blindness in Pigeons Columba Livia
    06:14

    A Method for Investigating Change Blindness in Pigeons Columba Livia

    Published on: September 7, 2018

    6.4K

    Area of Science:

    • Robotics and Control Systems
    • Artificial Intelligence
    • Game Theory

    Background:

    • Unmanned Aerial Vehicle (UAV) pursuit-evasion (PE) games present complex control challenges.
    • Existing methods often struggle with the coupled kinematics and dynamics inherent in these systems.
    • Efficiently computing Nash equilibrium (NE) solutions is crucial for developing robust UAV-PE strategies.

    Purpose of the Study:

    • To propose a novel hierarchical reinforcement learning (HRL) algorithm for UAV-PE game systems.
    • To integrate an alternative delay update (ADU) method for enhanced training efficiency.
    • To obtain approximate Nash equilibrium (NE) solutions for UAV-PE games with coupled kinematics and dynamics.

    Main Methods:

    • A hierarchical learning process involving zero-sum kinematics and optimal dynamics games.
    • Deep neural networks (NNs) to approximate policy and value functions at kinematic and dynamic levels.
    • The ADU method to stabilize training by fixing one player's strategy.

    Main Results:

    • Development of an HRL algorithm with ADU for UAV-PE systems.
    • Derivation of sufficient conditions for algorithm convergence and optimality.
    • Obtained overload inequalities ensuring dynamic state tracking via kinematic control input.

    Conclusions:

    • The proposed HRL algorithm with ADU is feasible and effective for UAV-PE games.
    • The method successfully computes approximate NE solutions for coupled systems.
    • Simulation results validate the algorithm's performance and training efficiency.