Related Experiment Video
Updated: May 22, 2025

06:25
A Real-Time Interactive System for Studying Confrontational Pursuit Behavior in Rodents
Published on: May 16, 2025
39
Multi-agent self-attention reinforcement learning for multi-USV hunting target
Summary
This study introduces a multi-agent reinforcement learning method using multi-head self-attention for multiple unmanned surface vehicles (multi-USV) to hunt targets. The novel approach enhances hunting success rates and algorithm convergence speed.
Area of Science:
- Robotics and Control Systems
- Artificial Intelligence
- Marine Engineering
Background:
- Cooperative target hunting with multiple unmanned surface vehicles (multi-USV) presents significant coordination challenges.
- Existing multi-agent reinforcement learning (MARL) algorithms may struggle with efficient information processing and rapid convergence in complex scenarios.
Purpose of the Study:
- To develop an advanced MARL method for multi-USV cooperative hunting.
- To enhance the efficiency and success rate of target acquisition by multi-USV systems.
Main Methods:
- Established kinematic, dynamic, and environmental models for USVs.
- Developed a differential game model for multi-USV target hunting, defining cooperative and competitive dynamics.
- Designed a MARL algorithm integrating the multi-head self-attention (MSA) mechanism to focus on critical information.
Main Results:
- The proposed MSA-based MARL algorithm demonstrated improved convergence speed by 17% compared to baseline MARL.
- Achieved an 8% increase in hunting success rate.
Conclusions:
- The MSA mechanism significantly enhances the performance of MARL algorithms in multi-USV cooperative hunting tasks.
- The developed method offers a more effective solution for autonomous multi-USV coordination and target interception.
Related Concept Videos
Multi-input and Multi-variable systems
93
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
In the absence...
93
Observational Learning
111
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
111
Masking and Demasking Agents
2.3K
EDTA titrations may necessitate masking and demasking agents to temporarily protect a particular metal ion in a mixture from the EDTA reaction. These agents facilitate the sequential analysis of the metal ions by forming stable complexes with some—but not all—metal ions during certain steps.
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
2.3K
Associative Learning
270
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
270
Avoidance Learning and Learned Helplessness
1.7K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.7K
Reinforcement
169
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
169

