持续时间强化学习:新的设计算法与理论见解和性能保证
概括
本研究引入了针对连续时间同源非线性系统的新的可刺激的整体强化学习 (EIRL) 算法. 这些方法可以增强激发,减少复杂性,从而在现实应用中提高控制性能和数据效率.
科学领域:
- 控制理论 控制理论
- 机器学习 机器学习
- 非线性系统是非线性系统.
背景情况:
- 持续时间强化学习 (CT-RL) 是有前途的,但缺乏实际应用.
- 适应动态编程 (ADP) 在理论上取得了成功,但实际演示有限.
- 现有的CT-RL方法面临着现实的控制挑战.
研究的目的:
- 引入新的可刺激的整体强化学习 (EIRL) 算法.
- 开发激发框架,以提高激发持续性 (PE) 和数值表现.
- 连续时间 (CT) 亲属非线性系统的地址控制.
主要方法:
- 使用经典的控制洞察力开发了一个新的激发框架.
- 将复杂的系统分成更小,更易于管理的子问题.
- 利用已知的同源非线性动力学来改善系统响应.
主要成果:
- 实现了良好的系统响应和增强数据效率.
- 通过将问题分解为子问题来证明复杂性降低.
- 提供了对融合,解决方案最佳性和闭环稳定性的保证.
结论:
- 对于现实的CT-RL控制问题,EIRL算法提供了可行的解决方案.
- 提出的方法确保了理论上的保证和实际的性能.
- 成功应用于控制不稳定的高超音速飞行器 (HSV).
更多相关视频
09:43A Fully Automated Rodent Conditioning Protocol for Sensorimotor Integration and Cognitive Control Experiments
Published on: April 15, 2014
10.6K
08:51Author Spotlight: Unveiling Neural Mechanisms Through Automated Evaluation of Motor Learning and Myelin Plasticity Studies Using the Erasmus Ladder
Published on: December 15, 2023
1.3K
相关概念视频
Reinforcement Schedules
144
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
144
Timing and Consequences on Behavior
90
In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective.
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
90
Reinforcement
202
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
202
Generalization, Discrimination, and Extinction
532
Generalization, discrimination, and extinction are key concepts in operant conditioning that influence how behaviors are learned and maintained.
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
532
Real-World Application of Classical Conditioning
550
Classical conditioning not only includes the initial pairing of stimuli but also extends to more complex forms, such as higher-order conditioning. Higher-order conditioning involves creating associations beyond the primary conditioned stimulus, resulting in a chain of conditioned responses.
Higher-order, or second-order, conditioning occurs when a neutral stimulus becomes associated with an already established conditioned stimulus through repeated pairings. For instance, if a dog has been...
Higher-order, or second-order, conditioning occurs when a neutral stimulus becomes associated with an already established conditioned stimulus through repeated pairings. For instance, if a dog has been...
550
Law of Effect
1.4K
B.F. Skinner, a prominent figure in behavioral psychology, introduced operant conditioning by emphasizing the role of consequences in shaping behavior. This theory builds upon the law of effect proposed by Edward Thorndike, which posits that behaviors followed by satisfying outcomes are likely to be repeated. In contrast, those followed by unsatisfying outcomes are less likely to recur.
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...
1.4K
