相关实验视频
Updated: Aug 19, 2025

09:45
New Variations for Strategy Set-shifting in the Rat
Published on: January 23, 2017
8.3K
通过无模型多代理强化学习掌握Stratego游戏
Julien Perolat1, Bart De Vylder1, Daniel Hennes1
1DeepMind Technologies Ltd., London, UK.
概括
我们开发了DeepNash, 一种掌握复杂的Stratego游戏的人工智能. 这种人工智能可以在Stratego中达到人类专家水平, 这是一款具有挑战性的不完美的信息游戏,
科学领域:
- 人工智能
- 游戏理论
- 强化学习
背景情况:
- 史特雷戈是一个标志性的棋盘游戏, 由于其战略深度和不完善信息的混合, 历史上挑战了人工智能.
- 掌握具有不完美的信息的游戏,比如克,以及象棋等战略规划,仍然是人工智能的重要研究领域.
研究的目的:
- 为了介绍DeepNash, 一个能够在人类专家水平上玩Stratego的AI代理.
- 展示一种新的方法来掌握不完美的信息游戏,使用强化学习而不是传统的搜索方法.
主要方法:
- DeepNash采用一种游戏理论,无模型的深度强化学习方法.
- 经纪人通过自主游戏自主学习Stratego战略,从头开始.
- 该方法特别省略了搜索算法的使用,使其与许多人工智能游戏技术有所区别.
主要成果:
- 迪普纳什在Stratego游戏中取得了人类专家水平的表现.
- 代理商在Stratego中超越了现有的最先进的人工智能方法.
- 在2022年,DeepNash在Gravon游戏平台上获得了前三名,与人类专家竞争.
结论:
- DeepNash代表了人工智能掌握复杂不完美的信息游戏的能力的重大进步.
- 对于Stratego来说,没有模型的深度强化学习方法是有效的.
- 这项工作突破了人工智能在战略游戏中的界限,
相关概念视频
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
97
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
97
Reinforcement
308
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
308
Reinforcement Schedules
228
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
228
Observational Learning
263
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
263
Multi-input and Multi-variable systems
139
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
In the absence...
139
Associative Learning
507
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
507

