Related Experiment Video
Updated: Sep 17, 2025

13:40
Combining Computer Game-Based Behavioural Experiments With High-Density EEG and Infrared Gaze Tracking
Published on: December 16, 2010
16.8K
The analysis of deep reinforcement learning for dynamic graphical games under artificial intelligence
Yuyang Yan1, Jiahui Li2, Cristina Zaggia3
1School of Education, Guangzhou University, Guangzhou, 510006, China.
Scientific Reports
|July 2, 2025
Summary
This study introduces a deep reinforcement learning (DRL) approach for autonomous decision-making in dynamic graphical games. The method enhances strategy optimization, improves accuracy, and reduces complexity for intelligent agents.
Area of Science:
- Artificial Intelligence
- Game Theory
- Control Systems
Background:
- Dynamic graphical games present complex challenges for autonomous decision-making.
- Traditional methods struggle with computational complexity and information exchange limitations.
- Optimizing strategies in these environments requires advanced learning techniques.
Purpose of the Study:
- To develop a deep reinforcement learning (DRL) framework for autonomous decision-making and strategy optimization in dynamic graphical games.
- To reduce computational complexity and information exchange through local performance metrics.
- To enable intelligent agents to make decisions using only local information.
Main Methods:
- An online iterative algorithm utilizing Deep Neural Networks (DNNs) and an Actor-Critic framework.
- Implementation of a distributed policy iteration mechanism for agent decision-making.
- Definition of local performance metrics to manage complexity.
Main Results:
- The DRL-based algorithm significantly improves decision accuracy and convergence speed.
- Demonstrated reduction in computational complexity and enhanced scalability.
- Effective performance in optimal control problems within dynamic graphical intelligent games.
Conclusions:
- The proposed DRL approach offers an effective solution for autonomous decision-making in complex graphical games.
- Local information-based distributed policy iteration enhances agent autonomy and efficiency.
- The method shows strong potential for real-world optimal control applications.
Related Concept Videos
Observational Learning
319
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
319
Reinforcement
353
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
353
Reinforcement Schedules
243
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
243
Dynamic Equilibrium
53.5K
A reversible chemical reaction represents a chemical process that proceeds in both forward (left to right) and reverse (right to left) directions. When the rates of the forward and reverse reactions are equal, the concentrations of the reactant and product species remain constant over time and the system is at equilibrium. A special double arrow is used to emphasize the reversible nature of the reaction. The relative concentrations of reactants and products in equilibrium systems vary greatly;...
53.5K
Avoidance Learning and Learned Helplessness
1.9K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.9K
Multi-input and Multi-variable systems
152
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
In the absence...
152

