Related Experiment Video
Updated: Jan 31, 2026

11:40
Quantitative Autonomic Testing
Published on: July 19, 2011
58.6K
Multi Pseudo Q-Learning-Based Deterministic Policy Gradient for Tracking Control of Autonomous Underwater Vehicles
IEEE Transactions on Neural Networks and Learning Systems
|January 4, 2019
Summary
This study introduces a novel Multi Pseudo Q-learning deterministic policy gradient (MPQ-DPG) algorithm for autonomous underwater vehicles (AUVs). The MPQ-DPG algorithm enhances trajectory tracking accuracy and learning stability in AUVs with unknown dynamics.
Area of Science:
- Robotics
- Control Systems
- Artificial Intelligence
Background:
- Underactuated autonomous underwater vehicles (AUVs) face challenges in trajectory tracking due to unknown dynamics and input constraints.
- Existing policy gradient methods with single actor-critic architectures often struggle with satisfactory tracking accuracy and stable learning.
Purpose of the Study:
- To develop an advanced algorithm for high-level trajectory tracking control accuracy and stable learning in underactuated AUVs.
- To address limitations of current methods by proposing a hybrid actors-critics architecture.
Main Methods:
- A hybrid actors-critics architecture is employed, training multiple actors and critics for deterministic policy and action-value function learning.
- Critics utilize an expected absolute Bellman error-based updating rule to select the worst critic for updates.
- Pseudo Q-learning and Multi Pseudo Q-learning (MPQ) are developed for continuous action spaces to reduce action-value function overestimation and stabilize learning.
- Deterministic policy gradient is applied to actors, with the final policy being an average of all actors to prevent drastic updates.
Main Results:
- The proposed MPQ-based deterministic policy gradient (MPQ-DPG) algorithm demonstrates effectiveness and generality in AUV trajectory tracking.
- The algorithm achieves high-level tracking control accuracy and stable learning across different reference trajectories.
- Increasing the number of actors and critics further enhances the performance of the MPQ-DPG algorithm.
Conclusions:
- The MPQ-DPG algorithm offers a significant improvement for trajectory tracking in underactuated AUVs with unknown dynamics.
- The hybrid actors-critics approach, combined with MPQ, effectively enhances control accuracy and learning stability.
- The findings suggest that the proposed method provides a robust solution for complex AUV control tasks.
Related Concept Videos
What is an Electrochemical Gradient?
127.8K
Adenosine triphosphate, or ATP, is considered the primary energy source in cells. However, energy can also be stored in the electrochemical gradient of an ion across the plasma membrane, which is determined by two factors: its chemical and electrical gradients.
The chemical gradient relies on differences in the abundance of a substance on the outside versus the inside of a cell and flows from areas of high to low ion concentration. In contrast, the electrical gradient revolves around an...
The chemical gradient relies on differences in the abundance of a substance on the outside versus the inside of a cell and flows from areas of high to low ion concentration. In contrast, the electrical gradient revolves around an...
127.8K
Autonomic Nervous System
12.8K
The autonomic nervous system (ANS) is a critical component of the peripheral nervous system, primarily responsible for regulating involuntary bodily functions and maintaining homeostasis. It functions in tandem with the central nervous system (CNS) to seamlessly coordinate various physiological processes without the need for conscious control.
The ANS comprises two main divisions: the sympathetic and parasympathetic divisions. These divisions function antagonistically to maintain a dynamic...
The ANS comprises two main divisions: the sympathetic and parasympathetic divisions. These divisions function antagonistically to maintain a dynamic...
12.8K
Autonomic Nervous System: Overview
7.5K
The human nervous system is divided into two main parts: the central nervous system (CNS) and the peripheral nervous system (PNS). The CNS is composed of the brain and spinal cord, while the PNS contains nerve cells, clusters of nerve cells, and the sensory receptors that are outside the CNS. The PNS has two types of nerve cells: sensory (afferent) and motor (efferent). Sensory cells send signals to the CNS from receptors, and motor cells carry signals from the CNS to organs, muscles, and...
7.5K
Disorders of the Autonomic Nervous System
1.6K
The autonomic nervous system (ANS) is an intricate network of nerves that controls functions such as the regulation of heart rate, digestion, and blood pressure regulation. When this system malfunctions, it can lead to various disorders that affect multiple bodily functions. One common feature of many autonomic disorders is the involvement of smooth blood vessels, which play a crucial role in regulating blood flow throughout the body.
Raynaud's disease, also known as Raynaud's...
Raynaud's disease, also known as Raynaud's...
1.6K
Avoidance Learning and Learned Helplessness
2.6K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
2.6K
Energy Line and Hydraulic Gradient Line
2.2K
Based on Bernoulli's equation, the energy line (EL) and hydraulic grade line (HGL) provide graphical representations of energy distribution in a fluid flow system. For steady, incompressible, inviscid flows, Bernoulli's equation is expressed as:
2.2K

