一般化政策改进算法与理论支持的样本重用
James Queeney1, Ioannis Ch Paschalidis2, Christos G Cassandras2
1Mitsubishi Electric Research Laboratories, Cambridge, MA 02139 USA. He performed the majority of this work while with the Division of Systems Engineering, Boston University, Boston, MA 02215 USA.
我们介绍了通用政策改进,这是一个新的类型的无模型深度强化学习算法. 这些算法平衡了性能保证与现实世界控制应用的数据效率.
科学领域:
- 人工智能的人工智能
- 机器学习 机器学习
- 控制系统 控制系统
背景情况:
- 无模型的深度强化学习 (RL) 对于数据驱动控制至关重要.
- 现有方法经常面临性能保证和数据效率之间的权衡.
- 现实世界的部署需要平衡这两个关键要求.
研究的目的:
- 开发一种新的类型的无模型深度RL算法.
- 在RL中解决性能保证与数据效率权衡的问题.
- 提高RL在实际控制场景中的适用性.
主要方法:
- 发展通用政策改进 (GPI) 算法.
- 结合政策上的方法保证与政策之外的样本重复使用效率.
- 对各种模拟控制任务进行了广泛的实验分析.
主要成果:
- 展示新的GPI算法的好处.
- 成功平衡性能保证和数据效率.
- 在广泛的模拟控制任务中进行验证.
结论:
- 拟议的GPI算法代表了无模型深度RL的重大进步.
- 这些算法为现实世界的控制问题提供了实际的解决方案.
- GPI算法有效地弥合了理论保障和实际数据效率之间的差距.
更多相关视频
11:53Spatial Multiobjective Optimization of Agricultural Conservation Practices using a SWAT Model and an Evolutionary Algorithm
Published on: December 9, 2012
11:53The Modular Design and Production of an Intelligent Robot Based on a Closed-Loop Control Strategy
Published on: October 14, 2017
相关概念视频
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Random Sampling Method
Systematic Sampling Method
Systematic sampling is one of the simplest methods...
Systematic Error: Methodological and Sampling Errors
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Trial and Error and Algorithm
