Related Experiment Video
Updated: Aug 12, 2025

A Networked Desktop Virtual Reality Setup for Decision Science and Navigation Experiments with Multiple Participants
Published on: August 26, 2018
Research on reinforcement learning-based safe decision-making methodology for multiple unmanned aerial vehicles
Longfei Yue1, Rennong Yang1, Ying Zhang1
1Air Traffic Control and Navigation College, Air Force Engineering University, Xi'an, China.
This study introduces a transfer-safe soft actor-critic (TSSAC) algorithm for safer and more efficient multi-unmanned aerial vehicle (multi-UAV) decision-making. The method enhances cooperative training and adaptability in dynamic environments.
Area of Science:
- Robotics and Artificial Intelligence
- Control Systems Engineering
- Autonomous Systems
Background:
- Multi-unmanned aerial vehicle (multi-UAV) systems offer significant advantages for complex tasks.
- Deep reinforcement learning (DRL) shows promise for multi-UAV decision-making but faces challenges in safety and training efficiency.
- Existing DRL methods require improvement for practical application in safety-critical multi-UAV operations.
Purpose of the Study:
- To develop a novel decision-making framework for multi-UAV systems that enhances both safety and training efficiency.
- To address the limitations of current DRL approaches in real-world multi-UAV deployments.
- To enable cooperative and adaptive control strategies for multi-UAVs in dynamic environments.
Main Methods:
- A transfer-safe soft actor-critic (TSSAC) algorithm is proposed for multi-UAV decision-making.
- Each UAV's decision-making is modeled as a constrained Markov decision process (CMDP) to incorporate safety constraints.
- The soft actor-critic-Lagrangian (SAC-Lagrangian) algorithm is integrated with a modified Lagrangian multiplier within the CMDP framework.
- Parameter-based transfer learning is employed to facilitate cooperative and efficient multi-UAV task training.
Main Results:
- Simulation experiments demonstrate that the TSSAC method significantly improves safety during multi-UAV operations.
- The proposed approach enhances the efficiency of training multi-UAV systems.
- The developed method enables UAVs to effectively adapt to dynamic and changing operational scenarios.
- The integration of CMDP and SAC-Lagrangian ensures that safety constraints are met while maximizing task performance.
Conclusions:
- The transfer-safe soft actor-critic (TSSAC) algorithm provides a viable solution for safe and efficient decision-making in multi-UAV systems.
- The proposed method effectively balances task performance with safety constraints through a modified CMDP formulation.
- Parameter-based transfer learning accelerates cooperative training and improves the adaptability of multi-UAVs.
- This research contributes to the practical deployment of DRL in complex, safety-critical autonomous systems.
More Related Videos
Related Concept Videos
Decision Making: Traditional Method
First, a specific claim about the population parameter is decided based on the research question and is stated in a simple form. Further, an opposing statement to this claim is also stated. These statements can act as null and alternative hypotheses, out of which a null hypothesis would be a...
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
Decision Making
Automatic decision-making is fast, intuitive, and relies on gut feelings...
Collisions in Multiple Dimensions: Problem Solving
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
Multi-input and Multi-variable systems
In the absence...
Response Surface Methodology
The process of RSM involves several key steps:

