Related Experiment Video
Updated: Jan 21, 2026

Structural Design and Manufacturing of a Cruiser Class Solar Vehicle
Published on: January 30, 2019
Critic-actor reinforcement learning for optimized cooperative formation of multi-nonholonomic wheeled Mobile vehicles
Lixia Liu1, Peiyong Duan1, Xiaoyu Liu1
1School of Mathematics and Statistics, Qilu University of Technology (Shandong Academy of Sciences), Jinan, 250353, China.
Abstract:
This study mainly focuses on an optimized leader-follower formation problem for a group of nonholonomic wheeled mobile vehicles (NWMVs). By systematically integrating the critic-actor reinforcement learning (RL) and adaptive neural network (NN), a distributed cooperative formation scheme for the multi-nonholonomic wheeled mobile vehicles (MNWMVs) consisting of a kinematic controller and a dynamic torque controller is proposed. The Hamilton-Jacobi-Bellman (HJB) equation, regarding the performance index function, possesses highly nonlinear and strongly coupled characteristics. It is challenging to solve the HJB equation to acquire the optimized formation protocol of multiple NWMVs, as it possesses under-actuated and nonholonomic Lagrange dynamic properties. Significantly, the key feature of the developed optimized formation tracking algorithm for MNWMVs is an adaptive identifier integrated into the critic-actor RL strategy. It effectively addresses the uncertainties associated with Lagrange dynamics. Furthermore, the developed optimized formation scheme is greatly simplified due to the RL training laws obtained from the negative gradient of a simple positive function. Finally, numerical simulations and physical experiments are performed to validate and demonstrate the theoretical results.
Related Concept Videos
Cooperative Allosteric Transitions
Cooperative Allosteric Transitions
Cooperative Allosteric Transitions
Cooperative Binding of Transcription Regulators
Actor-Observer Effect
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:

