适应性网络方法在强化学习中进行勘探-开发权衡
Mohammadamin Moradi1, Zheng-Meng Zhai1, Shirin Panahi1
1School of Electrical, Computer and Energy Engineering, Arizona State University, Tempe, Arizona 85287, USA.
Chaos (Woodbury, N.Y.)
|December 3, 2024
概括
这项研究引入了一种新的方法来平衡强化学习中的探索和利用,通过将其建模为非决定性的有限自动机. 这个框架优化了代理人的行动,以发现新的策略,并在未知的环境中最大限度地获得奖励.
科学领域:
- 人工智能的人工智能
- 机器学习 机器学习
- 计算理论 计算理论
背景情况:
- 强化学习 (RL) 代理人在平衡探索 (发现新策略) 和利用 (利用已知的策略) 中面临着一个关键的挑战.
- 对于代理人来说,有效的探索至关重要,以找到最佳的政策,并在不熟悉的环境中最大限度地获得长期回报.
- 剥削侧重于基于当前知识的即时收益,可能错过了优越的长期结果.
研究的目的:
- 开发一个系统的框架,在强化学习中平衡勘探和开发.
- 模拟强化学习过程作为一个非决定性的有限自动机.
- 优化代理人的行动,最大限度地发现新的,高奖励状态.
主要方法:
- 强化学习过程被概念化为马尔科夫决策过程,并以非决定性的有限自动机为模型.
- 自动机内的状态的一个子集被指定为代表对探索的偏好.
- 以混合整数编程 (MIP) 问题的形式制定了一个数学框架,通过优化代理行动来平衡探索和利用.
主要成果:
- 拟议的MIP配方提供了一种系统平衡勘探和开采的方法,从而产生最佳的权衡点.
- 在基准系统上的计算验证证明了框架的有效性.
- 展示了自动机作为一个具有不断变化的过渡概率的自适应网络,类似于复杂的动态网络.
结论:
- 开发的框架提供了一个基于原则的方法来管理强化学习中的勘探-开采困境.
- 强化学习自动机和自适应网络之间的联系为将网络理论应用于AI挑战开辟了新的途径.
- 这项研究促进了机器学习中的自适应系统的理解和应用.
相关概念视频
Social Exchange Theory
34.3K
We have discussed why we form relationships, what attracts us to others, and different types of love. But what determines whether we are satisfied with and stay in a relationship? One theory that provides an explanation is social exchange theory. According to social exchange theory, we act as naïve economists in keeping a tally of the ratio of costs and benefits of forming and maintaining a relationship with others (Rusbult & Van Lange, 2003).
34.3K
The Anchoring-and-Adjustment Heuristic
7.2K
In order to make good decisions, we use our knowledge and our reasoning. Often, this knowledge and reasoning is sound and solid. However, sometimes, we are swayed by biases or by others manipulating a situation. For example, let’s say you and three friends wanted to rent a house and had a combined target budget of $1,600. The realtor shows you only very run-down houses for $1,600 and then shows you a very nice house for $2,000. Might you ask each person to pay more in rent to get the...
7.2K
Instinctive Drift
187
Instinctive drift refers to the tendency of animals to revert to their innate behaviors despite repeated reinforcement. Breland and Breland demonstrated this concept in an experiment with a raccoon. The raccoon was trained to pick up two coins and place them in a container in exchange for food. Initially, the raccoon learned to associate the coins with food, making them a conditioned stimulus or a substitute for food. However, over time, the raccoon became less willing to put the coins into the...
187
Spare Receptors
3.5K
Some receptors remain unoccupied even when an agonist produces a maximal response. Such empty ones are called spare receptors. In presence of spare receptors the maximum effect of an agonist drug is achieved with fewer than 100% of the receptors being occupied. To determine the presence of spare receptors, scientists often compare the concentration of the drug needed to produce 50% of the maximum effect (EC50) with the concentration of the drug needed to occupy 50% of the receptors (Kd). If the...
3.5K


