Jove
Visualize
联系我们
JoVE
x logofacebook logolinkedin logoyoutube logo
关于 JoVE
概览领导团队博客JoVE 帮助中心
作者
出版流程编辑委员会范围与政策同行评审常见问题投稿
图书馆员
用户评价订阅访问资源图书馆顾问委员会常见问题
研究
JoVE JournalMethods CollectionsJoVE Encyclopedia of Experiments存档
教育
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab Manual教师资源中心教师网站
使用条款与条件
隐私政策
政策

相关概念视频

Reinforcement Schedules01:24

Reinforcement Schedules

459
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
459
Observational Learning01:12

Observational Learning

838
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
838
Reinforcement01:23

Reinforcement

839
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
839
Elaborative Rehearsals01:07

Elaborative Rehearsals

338
Elaborative rehearsal is a crucial cognitive strategy that strengthens information encoding in long-term memory by making meaningful connections between new data and pre-existing knowledge. This approach contrasts with maintenance rehearsal, which involves simple repetition without delving into the significance of the information. While maintenance rehearsal might temporarily keep information active in short-term memory, it is less effective for long-term retention.
The effectiveness of...
338
Associative Learning01:27

Associative Learning

1.2K
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
1.2K
Woodward–Hoffmann Selection Rules and Microscopic Reversibility01:34

Woodward–Hoffmann Selection Rules and Microscopic Reversibility

3.8K
Electrocyclic reactions, cycloadditions, and sigmatropic rearrangements are concerted pericyclic reactions that proceed via a cyclic transition state. These reactions are stereospecific and regioselective. The stereochemistry of the products depends on the symmetry characteristics of the interacting orbitals and the reaction conditions. Accordingly, pericyclic reactions are classified as either symmetry-allowed or symmetry-forbidden. Woodward and Hoffmann presented the selection criteria for...
3.8K

您也可能阅读

相关文章

通过共同作者、期刊和引用图与本文相关的文章。

排序
Same author

Identification of a 10-protein plasma signature for preclinical gout risk beyond serum urate: a robustness-first proteomic study in 33,147 normouricaemic individuals.

Journal of translational medicine·2026
Same author

Temperature Dependence of Fe Nucleation Behavior and Electrochemical Performance in Aqueous Fe Metal Batteries.

ACS applied materials & interfaces·2026
Same author

NEFL⁺NEFM⁺ myeloid-reprogrammed cells promote ccRCC progression through CX3CL1-CX3CR1-mediated Tregs chemotaxis.

Molecular and cellular biochemistry·2026
Same author

Strategies of high-accuracy memristor-based analogue computing in memory for artificial intelligence.

Nature materials·2026
Same author

Highly Cα-regio-, enantio- and diastereoselective Mukaiyama-type annulation of siloxyfurans: stereodivergent synthesis of multi-stereogenic tricyclic γ-lactones.

Chemical science·2026
Same author

Bifurcation of neural firing patterns driven by potassium dynamics and neuron-electrode geometry during high-frequency stimulation.

PLoS computational biology·2026

相关实验视频

Updated: Jan 16, 2026

Investigating Motor Skill Learning Processes with a Robotic Manipulandum
07:52

Investigating Motor Skill Learning Processes with a Robotic Manipulandum

Published on: February 12, 2017

9.2K

DT-HRL:通过重新想象的等级增强学习来掌握长序列操纵.

Junyang Zhang1, Yilin Zhang1, Honglin Sun1

  • 1Graduate School of Information, Production and Systems, Waseda University, Kitakyushu 808-0135, Japan.

Biomimetics (Basel, Switzerland)
|September 26, 2025
PubMed
概括

本研究介绍了一种使用机器人操纵器的决策转换器 (DT) 的等级强化学习 (HRL) 框架. 新方法在复杂的物流任务中增强了长期推理和概括.

科学领域:

  • 机器人技术 机器人技术 机器人技术
  • 人工智能的人工智能
  • 机器学习 机器学习

背景情况:

  • 物流中的机器人操纵者面临着多步骤任务,频繁切换和长期依赖的挑战.
  • 现有的方法在动态环境中难以处理复杂的,连续的决策.
  • 人类运动控制为层次任务执行提供了一个模型.

研究的目的:

  • 为机器人操纵者提出一个新的等级强化学习 (HRL) 框架.
  • 改善物流中的长期推理,概括和任务执行.
  • 将决策转换器 (DT) 功能与层次控制结构集成.

主要方法:

  • 开发了一个多任务目标条件决策转换器 (MTGC-DT) 框架.
  • 一个高级别的政策模型将马尔科夫决策过程作为一个序列建模任务.
  • 一个低级策略使用参数化的动作原始体进行物理执行.
  • 引入了路径效率损失 (PEL) 校正和可学习的原始技能库.

主要成果:

  • 基于决策变压器的层次增强学习 (DT-HRL) 与基线相比,成功率超过10%.
  • 在物流任务中,DT-HRL的平均报酬超过了8%.
关键词:
决策变换器 决策变换器层次化的强化学习学习.长时间的任务顺序.机器人操纵是一种机器人操纵.

更多相关视频

The Double-H Maze: A Robust Behavioral Test for Learning and Memory in Rodents
09:01

The Double-H Maze: A Robust Behavioral Test for Learning and Memory in Rodents

Published on: July 8, 2015

13.1K
Utilizing a Reconfigurable Maze System to Enhance the Reproducibility of Spatial Navigation Tests in Rodents
04:41

Utilizing a Reconfigurable Maze System to Enhance the Reproducibility of Spatial Navigation Tests in Rodents

Published on: December 2, 2022

3.3K

相关实验视频

Last Updated: Jan 16, 2026

Investigating Motor Skill Learning Processes with a Robotic Manipulandum
07:52

Investigating Motor Skill Learning Processes with a Robotic Manipulandum

Published on: February 12, 2017

9.2K
The Double-H Maze: A Robust Behavioral Test for Learning and Memory in Rodents
09:01

The Double-H Maze: A Robust Behavioral Test for Learning and Memory in Rodents

Published on: July 8, 2015

13.1K
Utilizing a Reconfigurable Maze System to Enhance the Reproducibility of Spatial Navigation Tests in Rodents
04:41

Utilizing a Reconfigurable Maze System to Enhance the Reproducibility of Spatial Navigation Tests in Rodents

Published on: December 2, 2022

3.3K
  • 废弃实验显示,正常化得分增加了2%以上.
  • 结论:

    • 拟议的DT-HRL框架有效地解决了长期的依赖关系,并改善了机器人操纵的泛化.
    • 将决策转换器与HRL集成为物流中复杂任务自动化提供了一个有希望的方向.
    • 该框架的模块化设计与参数化技能增强了可重复使用性和适应性.