神经Q学习用于离散时间非线性零和游戏,可调节的收率
Yuan Wang1, Ding Wang1, Mingming Zhao1
1Faculty of Information Technology, Beijing University of Technology, Beijing 100124, China; Beijing Key Laboratory of Computational Intelligence and Intelligent System, Beijing University of Technology, Beijing 100124, China; Beijing Institute of Artificial Intelligence, Beijing University of Technology, Beijing 100124, China; Beijing Laboratory of Smart Environmental Protection, Beijing University of Technology, Beijing 100124, China.
本研究引入了一个可调整的Q学习方案,以加快使用神经网络解决非线性零和游戏的速度. 开发的算法确保了融合,并在模拟中展示了出色的性能.
科学领域:
- 控制理论 控制理论
- 人工智能的人工智能
- 游戏理论 游戏理论
背景情况:
- 零和游戏带来复杂的控制挑战.
- 实际应用需要无模型方法.
- 加快代学习算法对于效率至关重要.
研究的目的:
- 为离散时间非线性零和游戏开发可调节的Q学习方案.
- 为了提高Q函数代的收率.
- 通过神经网络实现无模型的跟踪控制.
主要方法:
- 对代Q函数的单调性和收性的分析.
- 神经网络的集成,以实现无模型控制.
- 设计了两个加速的Q学习算法,并保证了趋同.
主要成果:
- 一个可调整的Q学习方案加速了融合.
- 两个算法确保与自适应加速阶段或可调节的放松因子的融合.
- 模拟证实了算法在实际场景中的有效性.
结论:
- 提出的可调整的Q学习方案有效地解决了非线性零和游戏.
- 神经网络集成促进了无模型的跟踪控制.
- 开发的算法为Q学习提供了加速和保证的融合.
相关概念视频
Current Growth And Decay In RL Circuits
BIBO stability of continuous and discrete -time systems
To determine the BIBO stability, the convolution integral is utilized when a bounded continuous-time input is applied to a Linear Time-Invariant (LTI) system....
Feedback control systems
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
Linear time-invariant Systems
The input-output behavior of an LTI system can be fully defined by its response to an impulsive excitation at its input. Once this impulse response is known, the system's reaction to any other input can be...
Linear Approximation in Time Domain
For a simple pendulum with a mass evenly distributed along its length and the center of mass located at half the pendulum's length,...
Region of Convergence


