相关实验视频
Updated: Jan 11, 2026

06:48
The HoneyComb Paradigm for Research on Collective Human Behavior
Published on: January 19, 2019
9.8K
一个统一和高效的培训框架,用于开放式的非过渡性游戏
Shaokang Dong1, Chao Li2, Shangdong Yang2
1School of Computer and Electronic Information, Nanjing Normal University, Nanjing, China; State Key Laboratory for Novel Software Technology, Nanjing University, China.
概括
本研究介绍了一种统一和有效的游戏 (UEP) 框架,用于解决复杂游戏中的纳什平衡 (NE). UEP结合了自我游戏 (SP) 和政策空间响应预言 (PSRO) 以实现更快,更强大的政策生成.
科学领域:
- 游戏理论 游戏理论
- 人工智能的人工智能
- 强化学习是一种强化学习.
背景情况:
- 自动游戏 (SP) 和政策空间响应预言 (PSRO) 是在游戏中寻找纳什平衡 (NE) 的关键框架.
- SP在过渡性游戏中表现出色,但在非过渡性游戏中扎.
- PSRO处理非过渡性游戏,但由于从头开始重新培训政策,其效率低下.
研究的目的:
- 开发一个新的框架,统一和高效的游戏 (UEP),合并了SP和PSRO的优势.
- 在具有热启动能力的开放式,非过渡性的游戏中实现高效的NE解决方案.
- 通过平衡政策准确性和多样性来提高培训效率.
主要方法:
- 引入统一和有效的游戏 (UEP) 框架.
- 开发统一的规范化距离度量,以平衡SP精度和PSRO多样性.
- 对UEP趋同到NE的理论分析.
- 在各种开放式,非过渡性的游戏中进行实证验证.
主要成果:
- UEP框架有效地结合了SP和PSRO的优势.
- 拟议的规范化距离指标提高了培训效率.
- 建立了UEP到NE的理论融合.
- 经验结果表明,UEP在接近NE和产生强有力的政策方面表现优越.
结论:
- 在具有挑战性的游戏环境中,UEP为解决NE提供了显著的进步.
- 该框架为现有的SP和PSRO方法提供了一个更有效和更强大的替代方案.
- UEP显示了复杂,开放式游戏中的应用的巨大潜力.
相关概念视频
Collisions in Multiple Dimensions: Problem Solving
5.2K
In multiple dimensions, the conservation of momentum applies in each direction independently. Hence, to solve collisions in multiple dimensions, we should write down the momentum conservation in each direction separately. To help understand collisions in multiple dimensions, consider an example.
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
5.2K
Observational Learning
804
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
804
Collisions in Multiple Dimensions: Introduction
6.5K
It is far more common for collisions to occur in two dimensions; that is, the initial velocity vectors are neither parallel nor antiparallel to each other. Let's see what complications arise from this. The first idea is that momentum is a vector. Like all vectors, it can be expressed as a sum of perpendicular components (usually, though not always, an x-component and a y-component, and a z-component if necessary). Thus, when the statement of conservation of momentum is written for a...
6.5K
Statically Indeterminate Problem Solving
675
Statically indeterminate problems are those where statics alone can not determine the internal forces or reactions. Consider a structure comprising two cylindrical rods made of steel and brass. These rods are joined at point B and restrained by rigid supports at points A and C. Now, the reactions at points A and C and the deflection at point B are to be determined. This rod structure is classified as statically indeterminate as the structure has more supports than are necessary for maintaining...
675
Generalization, Discrimination, and Extinction
1.3K
Generalization, discrimination, and extinction are key concepts in operant conditioning that influence how behaviors are learned and maintained.
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
1.3K
Multi-input and Multi-variable systems
378
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence of...
In the absence of...
378