相关实验视频
Updated: Jan 4, 2026

06:04
Study Motor Skill Learning by Single-pellet Reaching Tasks in Mice
Published on: March 4, 2014
22.0K
在StarCraft II中使用多代理增强学习的大师级别
Oriol Vinyals1, Igor Babuschkin2, Wojciech M Czarnecki2
1DeepMind, London, UK. vinyals@google.com.
Nature
|November 1, 2019
概括
阿尔法星,一个人工智能代理,通过使用多代理增强学习在"星际战斗2"中达到大师级别. 这种人工智能在复杂的现实世界战略游戏中表现出先进的能力.
科学领域:
- 人工智能
- 计算游戏理论
- 多个代理系统
背景情况:
- 星际飞船具有复杂的多代理挑战与现实应用相关.
- 在"星际战斗"之前,人工智能代理人依赖于游戏简化或超人能力.
- 之前没有任何人工智能能够与人类星际飞船玩家的技能相匹配.
研究的目的:
- 在复杂的实时策略游戏"星际争夺者2"中,
- 利用适用于其他具有挑战性的领域的通用学习方法.
主要方法:
- 采用多个代理强化学习算法.
- 使用各种各样的深度神经网络与适应性策略和反策略.
- 在"星际战斗2"中使用人类和代理游戏的数据进行训练.
主要成果:
- 这位阿尔法星特工在所有三款星际争夺战中都取得了大师级别.
- 阿尔法星的表现超过了人类玩家的99.8%.
- 在复杂的战略游戏中展示了人工智能能力的显著进步.
结论:
- 一般目的的学习方法,特别是多代理强化学习,可以在复杂的战略环境中实现人类水平的表现,如StarCraft II.
- 它代表了人工智能研发实时战略游戏的重要里程碑.
- 这种方法可以扩展到需要复杂决策和协调的其他领域.
相关概念视频
Reinforcement
765
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
765
Observational Learning
769
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
769
Reinforcement Schedules
414
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
414
Avoidance Learning and Learned Helplessness
2.4K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
2.4K
Multi-input and Multi-variable systems
358
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence of...
In the absence of...
358

