DT-HRL:通过重新想象的等级增强学习来掌握长序列操纵
Junyang Zhang1, Yilin Zhang1, Honglin Sun1
1Graduate School of Information, Production and Systems, Waseda University, Kitakyushu 808-0135, Japan.
Biomimetics (Basel, Switzerland)
|September 26, 2025
概括
本研究介绍了一种使用机器人操纵器的决策转换器 (DT) 的等级强化学习 (HRL) 框架. 新方法在复杂的物流任务中增强了长期推理和概括.
科学领域:
- 机器人技术 机器人技术 机器人技术
- 人工智能的人工智能
- 机器学习 机器学习
背景情况:
- 物流中的机器人操纵者面临着多步骤任务,频繁切换和长期依赖的挑战.
- 现有的方法在动态环境中难以处理复杂的,连续的决策.
- 人类运动控制为层次任务执行提供了一个模型.
研究的目的:
- 为机器人操纵者提出一个新的等级强化学习 (HRL) 框架.
- 改善物流中的长期推理,概括和任务执行.
- 将决策转换器 (DT) 功能与层次控制结构集成.
主要方法:
- 开发了一个多任务目标条件决策转换器 (MTGC-DT) 框架.
- 一个高级别的政策模型将马尔科夫决策过程作为一个序列建模任务.
- 一个低级策略使用参数化的动作原始体进行物理执行.
- 引入了路径效率损失 (PEL) 校正和可学习的原始技能库.
主要成果:
- 基于决策变压器的层次增强学习 (DT-HRL) 与基线相比,成功率超过10%.
- 在物流任务中,DT-HRL的平均报酬超过了8%.
- 废弃实验显示,正常化得分增加了2%以上.
结论:
- 拟议的DT-HRL框架有效地解决了长期的依赖关系,并改善了机器人操纵的泛化.
- 将决策转换器与HRL集成为物流中复杂任务自动化提供了一个有希望的方向.
- 该框架的模块化设计与参数化技能增强了可重复使用性和适应性.
相关概念视频
Reinforcement Schedules
459
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
459
Observational Learning
838
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
838
Reinforcement
839
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
839
Elaborative Rehearsals
338
Elaborative rehearsal is a crucial cognitive strategy that strengthens information encoding in long-term memory by making meaningful connections between new data and pre-existing knowledge. This approach contrasts with maintenance rehearsal, which involves simple repetition without delving into the significance of the information. While maintenance rehearsal might temporarily keep information active in short-term memory, it is less effective for long-term retention.
The effectiveness of...
The effectiveness of...
338
Associative Learning
1.2K
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
1.2K
Woodward–Hoffmann Selection Rules and Microscopic Reversibility
3.8K
Electrocyclic reactions, cycloadditions, and sigmatropic rearrangements are concerted pericyclic reactions that proceed via a cyclic transition state. These reactions are stereospecific and regioselective. The stereochemistry of the products depends on the symmetry characteristics of the interacting orbitals and the reaction conditions. Accordingly, pericyclic reactions are classified as either symmetry-allowed or symmetry-forbidden. Woodward and Hoffmann presented the selection criteria for...
3.8K


