Related Experiment Video
Updated: Jul 8, 2025

A Real-Time Interactive System for Studying Confrontational Pursuit Behavior in Rodents
Published on: May 16, 2025
Reinforcement learning-based formation-surrounding control for multiple quadrotor UAVs pursuit-evasion games
1Logistics Engineering College, Shanghai Maritime University, Shanghai 201306, China.
This study introduces a reinforcement learning (RL) control method for multiple quadrotor unmanned aerial vehicles (UAVs) in pursuit-evasion games. The method ensures pursuers equally surround evaders despite external disturbances.
Area of Science:
- Robotics
- Control Systems
- Artificial Intelligence
Background:
- Multiple quadrotor unmanned aerial vehicles (UAVs) are increasingly used in complex scenarios.
- Pursuit-evasion (PE) games present significant control challenges, especially with multiple agents and external disturbances.
- Achieving stable formations and coordinated maneuvers in multi-UAV systems requires advanced control strategies.
Purpose of the Study:
- To develop a reinforcement learning (RL)-based formation-surrounding control method for multiple UAVs in pursuit-evasion (MPE) games.
- To address the challenge of pursuers equally surrounding evaders while maintaining formation integrity under external disturbances.
- To ensure the stability of quadrotor UAVs' position and attitude tracking error subsystems.
Main Methods:
- Proposed a novel control strategy combining feedforward control and reinforcement learning (RL).
- Developed two new cost functions tailored for quadrotor UAVs facing external disturbances.
- Implemented critic-only neural network (NN) weight update laws for optimal cost function estimation.
- Utilized RL to solve Hamilton-Jacobi-Isaacs (HJI) equations for achieving Nash equilibrium.
- Applied Euler's formula to prove the property of equal surrounding.
Main Results:
- Demonstrated the stability of the tracking error subsystem using RL-based control schemes.
- Successfully achieved Nash equilibrium for the multiple quadrotor UAV system.
- Proved the property of equal surrounding for pursuers around evaders for the first time.
- Numerical simulations confirmed the effectiveness and superior performance of the proposed method.
Conclusions:
- The proposed RL-based formation-surrounding control method is effective for multi-UAV pursuit-evasion games.
- The method ensures stable control and achieves the desired equal surrounding formation under external disturbances.
- This research advances multi-agent control systems by integrating RL for complex coordinated behaviors.
Related Concept Videos
Collisions in Multiple Dimensions: Problem Solving
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Collisions in Multiple Dimensions: Introduction
Three-Dimensional Force System:Problem Solving
To solve a three-dimensional force system, first resolve each force into its respective scalar components. Do this using...
Observational Learning
Open and closed-loop control systems
An open-loop control system operates without feedback from the output. It consists of two primary elements: the controller and the controlled process. The controller receives an input signal...

