Related Experiment Video
Updated: Nov 24, 2025

08:18
WheelCon: A Wheel Control-Based Gaming Platform for Studying Human Sensorimotor Control
Published on: August 15, 2020
5.2K
Multiplayer Stackelberg-Nash Game for Nonlinear System via Value Iteration-Based Integral Reinforcement Learning.
IEEE Transactions on Neural Networks and Learning Systems
|December 22, 2020
Summary
This study introduces a novel reinforcement learning algorithm for multiplayer Stackelberg-Nash games in nonlinear systems. The algorithm efficiently finds equilibrium strategies using partial system dynamics, verified by simulations.
Area of Science:
- Game Theory
- Control Theory
- Reinforcement Learning
Background:
- Multiplayer Stackelberg-Nash games (SNG) involve hierarchical decision-making with one leader and multiple followers.
- Analyzing nonlinear dynamical systems in SNGs presents computational challenges for finding equilibrium points.
Purpose of the Study:
- To develop a novel algorithm for solving multiplayer SNGs in nonlinear dynamical systems.
- To address the analytical difficulties in calculating equilibrium strategies.
- To ensure the admissibility and convergence of the proposed solution.
Main Methods:
- A two-level value iteration-based integral reinforcement learning (VI-IRL) algorithm was developed.
- The algorithm utilizes partial information of system dynamics and neural networks for function approximation.
- Least-squares methods are employed for updating weights.
Main Results:
- The VI-IRL algorithm converges asymptotically to the Stackelberg-Nash equilibrium strategies under weak coupling conditions.
- Effective termination criteria were introduced to guarantee policy admissibility after finite iterations.
- The algorithm's effectiveness was validated through two simulation examples.
Conclusions:
- The proposed VI-IRL algorithm offers an effective method for solving complex multiplayer SNGs in nonlinear systems.
- The approach overcomes analytical limitations by using reinforcement learning and partial system information.
- The method provides a computationally feasible way to determine equilibrium strategies.
Related Concept Videos
Multi-input and Multi-variable systems
267
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence of...
In the absence of...
267
Statically Indeterminate Problem Solving
587
Statically indeterminate problems are those where statics alone can not determine the internal forces or reactions. Consider a structure comprising two cylindrical rods made of steel and brass. These rods are joined at point B and restrained by rigid supports at points A and C. Now, the reactions at points A and C and the deflection at point B are to be determined. This rod structure is classified as statically indeterminate as the structure has more supports than are necessary for maintaining...
587
Reinforcement
617
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
617
Application of Nonlinear Inequalities
76
A nonlinear inequality describes a comparison involving an expression that curves or behaves more complexly than a straight line. These inequalities often appear in forms that include squares, products, or variables in the denominator.To solve such an inequality, one starts by rewriting it so that zero appears on one side. For example, the inequality: can be factored as: This form makes it easier to identify the values that cause the expression to equal zero. In this case, the key values...
76
Reinforcement Schedules
331
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
331
Introduction to Nonlinear Inequalities
70
Linear and nonlinear inequalities are fundamental for analyzing variable relationships and identifying ranges satisfying specific conditions. A linear inequality involves variables raised only to the first power, resulting in a straight-line graph. This line partitions the coordinate plane into two distinct regions: one that satisfies the inequality and one that does not. Each region represents a set of solutions where the linear relationship holds true under the specified constraint.Nonlinear...
70

