折扣规范化的意想不到的后果:提高规范化的确定性等价性加强学习学习
Sarah Rathnam1, Sonali Parbhoo2, Weiwei Pan1
1Harvard University, School of Engineering and Applied Sciences, Cambridge, MA USA.
概括
马尔科夫决策流程 (MDP) 中的折扣规范化可能会导致数据不均的政策估计不佳. 这项研究揭示了折扣规范化像状态-动作前期一样,为状态-动作特定的规范化参数提出了一种新方法.
科学领域:
- 人工智能的人工智能
- 机器学习 机器学习
- 强化学习是一种强化学习.
背景情况:
- 折扣规范化通常用于强化学习,以简化从稀疏或杂数据的政策估计.
- 它通常被理解为在马尔科夫决策过程 (MDP) 中减轻延迟效应的重视.
研究的目的:
- 揭示对折扣规范化的另一个观点,强调其意想不到的后果.
- 为了证明折扣规范化相当于过渡矩阵上的前期.
- 提出一种新的方法来设定国家特定的规范化参数.
主要方法:
- 开发了一个等价定理,表明折扣规范化相当于过渡矩阵上的先验.
- 导出了国家行动特定规范化参数的明确公式.
- 用模拟和医疗癌症模拟器实证评估了与标准折扣规范化对比的拟议方法.
主要成果:
- 证明了折扣正规化作为一个先验的功能,在拥有更多数据的状态-动作对上有更强的正规化.
- 展示了当从不均的数据集估计过渡矩阵时,这会导致低于最佳的性能.
- 拟议的国家具体行动方法显著提高了政策估计的准确性.
结论:
- 折扣规范化具有意想不到的后果,作为数据依赖的先验.
- 一种特定于国家行动的规范化方法弥补了全球折扣规范化的缺点.
- 拟议的方法为强化学习提供了更好的政策估计,特别是在数据不均的情况下.
更多相关视频
05:22Dissociation of the Confounding Influences of Expectancy and Integrative Difficulty Residing in Anomalous Sentences in Event-related Potential Studies
Published on: May 9, 2019
5.4K
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
618
相关概念视频
Propagation of Uncertainty from Random Error
725
An experiment often consists of more than a single step. In this case, measurements at each step give rise to uncertainty. Because the measurements occur in successive steps, the uncertainty in one step necessarily contributes to that in the subsequent step. As we perform statistical analysis on these types of experiments, we must learn to account for the propagation of uncertainty from one step to the next. The propagation of uncertainty depends on the type of arithmetic operation performed on...
725
Generalization, Discrimination, and Extinction
609
Generalization, discrimination, and extinction are key concepts in operant conditioning that influence how behaviors are learned and maintained.
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
609
Propagation of Uncertainty from Systematic Error
553
The atomic mass of an element varies due to the relative ratio of its isotopes. A sample's relative proportion of oxygen isotopes influences its average atomic mass. For instance, if we were to measure the atomic mass of oxygen from a sample, the mass would be a weighted average of the isotopic masses of oxygen in that sample. Since a single sample is not likely to perfectly reflect the true atomic mass of oxygen for all the molecules of oxygen on Earth, the mass we obtain from this...
553
Constraints and Statical Determinacy
634
In structural engineering, the equilibrium of a system is not only determined by its equations of equilibrium but also with the help of constraints. Constraints refer to restrictions on the motion of a system. The proper combinations of constraints can minimize the total number of constraints needed to maintain a system in mechanical equilibrium. When this happens, the system is said to be statically determinate. For such systems, the unknown reaction supports can be estimated using equilibrium...
634
Reinforcement
274
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
274
Hindsight Biases
3.4K
Hindsight bias leads you to believe that the event you just experienced was predictable, even though it really wasn’t. In other words, you knew all along that things would turn out the way they did. Can you relate this to the phrase "Hindsight is 20/20" now?
3.4K
