在人工智能下的动态图形游戏深度强化学习的分析
Yuyang Yan1, Jiahui Li2, Cristina Zaggia3
1School of Education, Guangzhou University, Guangzhou, 510006, China.
Scientific reports
|July 2, 2025
概括
本研究引入了深度强化学习 (DRL) 方法,用于动态图形游戏中的自主决策. 该方法增强了战略优化,提高了准确性,并减少了智能代理的复杂性.
科学领域:
- 人工智能的人工智能
- 游戏理论 游戏理论
- 控制系统 控制系统
背景情况:
- 动态图形游戏对自主决策提出了复杂的挑战.
- 传统的方法与计算复杂性和信息交换局限性作斗争.
- 在这些环境中优化策略需要先进的学习技术.
研究的目的:
- 在动态图形游戏中开发深度强化学习 (DRL) 框架,用于自主决策和战略优化.
- 通过本地性能指标来减少计算复杂性和信息交换.
- 使智能代理能够仅使用本地信息做出决策.
主要方法:
- 一个在线代算法,利用深度神经网络 (DNN) 和一个Actor-Critic框架.
- 实施一个分布式的政策代机制,用于代理决策.
- 定义本地绩效指标以管理复杂性.
主要成果:
- 基于DRL的算法显著提高了决策准确性和融合速度.
- 证明了计算复杂性的降低和增强的可扩展性.
- 在动态图形智能游戏中的最佳控制问题中的有效性能.
结论:
- 拟议的DRL方法为复杂的图形游戏中的自主决策提供了有效的解决方案.
- 基于本地信息的分布式政策代增强了代理人的自主性和效率.
- 该方法显示了对现实世界最佳控制应用的巨大潜力.
相关概念视频
Observational Learning
319
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
319
Reinforcement
353
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
353
Reinforcement Schedules
243
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
243
Dynamic Equilibrium
53.5K
A reversible chemical reaction represents a chemical process that proceeds in both forward (left to right) and reverse (right to left) directions. When the rates of the forward and reverse reactions are equal, the concentrations of the reactant and product species remain constant over time and the system is at equilibrium. A special double arrow is used to emphasize the reversible nature of the reaction. The relative concentrations of reactants and products in equilibrium systems vary greatly;...
53.5K
Avoidance Learning and Learned Helplessness
1.9K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.9K
Multi-input and Multi-variable systems
152
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
In the absence...
152


