Jove
Visualize
联系我们
JoVE
x logofacebook logolinkedin logoyoutube logo
关于 JoVE
概览领导团队博客JoVE 帮助中心
作者
出版流程编辑委员会范围与政策同行评审常见问题投稿
图书馆员
用户评价订阅访问资源图书馆顾问委员会常见问题
研究
JoVE JournalMethods CollectionsJoVE Encyclopedia of Experiments存档
教育
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab Manual教师资源中心教师网站
使用条款与条件
隐私政策
政策

相关概念视频

Role of Shaping in Operant Conditioning01:19

Role of Shaping in Operant Conditioning

251
Shaping is a technique used in operant conditioning to train complex behaviors by rewarding successive approximations toward the target behavior. This method is necessary because organisms are unlikely to perform complex behaviors spontaneously. Instead, shaping breaks down the desired behavior into small, manageable steps.
The steps involved in shaping begin with reinforcing any response that resembles the desired behavior. For example, parents might praise a child for picking up one toy. As...
251
Reinforcement Schedules01:24

Reinforcement Schedules

127
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
127
Graded Potential01:19

Graded Potential

3.5K
Graded potentials are localized fluctuations in the cell membrane's electrical charge, commonly found in the dendrites of neurons. The magnitude of these potential changes depends on the strength of the initiating stimulus. In a membrane at its resting potential, a graded potential signifies a voltage shift either above -70 mV or below -70 mV.
Graded potentials fall into two categories: depolarizing and hyperpolarizing. Depolarizing graded potentials typically occur when sodium (Na+) or...
3.5K
Reinforcement01:23

Reinforcement

176
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
176
Primary and Secondary Reinforcers01:23

Primary and Secondary Reinforcers

162
In psychology, reinforcement is a key concept in behavior modification. B.F. Skinner demonstrated this with his experiments involving rats in what is known as a Skinner box. The rats learned to press a lever to receive food, a primary reinforcer that fulfilled their innate need for nourishment.
Effective reinforcers for humans vary depending on the individual and the context. Primary reinforcers, such as food, water, sleep, shelter, and pleasure, have inherent value and satisfy basic biological...
162

您也可能阅读

相关文章

通过共同作者、期刊和引用图与本文相关的文章。

排序
Same author

Multi-structure segmentation in CBCT volumes: The ToothFairy2 challenge.

Medical image analysis·2026
Same author

Automated Immunophenotyping Assessment for Diagnosing Childhood Acute Leukemia using Set-Transformers.

Annual International Conference of the IEEE Engineering in Medicine and Biology Society. IEEE Engineering in Medicine and Biology Society. Annual International Conference·2025
Same author

The impact of multicentric datasets for the automated tumor delineation in primary prostate cancer using convolutional neural networks on <sup>18</sup>F-PSMA-1007 PET.

Radiation oncology (London, England)·2024
Same author

Deep learning based automated delineation of the intraprostatic gross tumour volume in PSMA-PET for patients with primary prostate cancer.

Radiotherapy and oncology : journal of the European Society for Therapeutic Radiology and Oncology·2023
Same author

Investigation and benchmarking of U-Nets on prostate segmentation tasks.

Computerized medical imaging and graphics : the official journal of the Computerized Medical Imaging Society·2023
Same author

Optimized Detection of High-Dimensional Entanglement.

Physical review letters·2021

相关实验视频

Updated: May 26, 2025

Automated Visual Cognitive Tasks for Recording Neural Activity Using a Floor Projection Maze
11:15

Automated Visual Cognitive Tasks for Recording Neural Activity Using a Floor Projection Maze

Published on: February 20, 2014

13.0K

HPRS:基于潜在的等级奖励,根据任务规范来塑造奖励.

Luigi Berducci1, Edgar A Aguilar2, Dejan Ničković2

  • 1Cyber-Physical Systems Group, Computer Engineering, TU Wien, Vienna, Austria.

Frontiers in robotics and AI
|February 25, 2025
PubMed
概括

本研究介绍了机器人强化学习的层次,基于潜力的奖励塑造 (HPRS). 高PRS有效地平衡了多个要求,改善了政策性能,并使无的sim-to-real转移成为可能.

关键词:
正式的规格,正式的规格.强化学习是一种强化学习.奖励塑造方式奖励塑造方式机器人学习机器人学习机器人技术 机器人工程 机器人工程

更多相关视频

Investigating Motor Skill Learning Processes with a Robotic Manipulandum
07:52

Investigating Motor Skill Learning Processes with a Robotic Manipulandum

Published on: February 12, 2017

8.7K
Study Motor Skill Learning by Single-pellet Reaching Tasks in Mice
06:04

Study Motor Skill Learning by Single-pellet Reaching Tasks in Mice

Published on: March 4, 2014

20.8K

相关实验视频

Last Updated: May 26, 2025

Automated Visual Cognitive Tasks for Recording Neural Activity Using a Floor Projection Maze
11:15

Automated Visual Cognitive Tasks for Recording Neural Activity Using a Floor Projection Maze

Published on: February 20, 2014

13.0K
Investigating Motor Skill Learning Processes with a Robotic Manipulandum
07:52

Investigating Motor Skill Learning Processes with a Robotic Manipulandum

Published on: February 12, 2017

8.7K
Study Motor Skill Learning by Single-pellet Reaching Tasks in Mice
06:04

Study Motor Skill Learning by Single-pellet Reaching Tasks in Mice

Published on: March 4, 2014

20.8K

科学领域:

  • 机器人技术 机器人技术 机器人技术
  • 人工智能的人工智能
  • 控制理论 控制理论

背景情况:

  • 机器人政策合成的强化学习 (RL) 在很大程度上取决于奖励信号.
  • 目前的方法很难产生有效的奖励信号,以满足各种各样的高层要求.
  • 从正式要求定义自动奖励是一个活跃的研究领域,具有现有的局限性.

研究的目的:

  • 引入一种自动化方法来生成准确反映层次任务要求的奖励信号.
  • 开发一种新的方法,分层的,基于潜力的奖励塑造 (HPRS),用于创建有效和多目标的奖励功能.
  • 证明HPRS在制造满足复杂安全,目标和舒适性要求的政策方面的能力.

主要方法:

  • 将任务定义为安全,目标和舒适性要求的部分顺序集合.
  • 将这些要求自动转化为等级奖励结构,奖励是彼此的函数.
  • 采用基于潜力的奖励塑造,将稀疏的奖励转化为密集的奖励,同时保持政策的最佳性.
  • 在八个机器人基准和两个模拟到真实的F1TENTH车辆应用上进行实验.

主要成果:

  • 在各种机器人基准中,HPRS成功地生成了满足各种机器人基准的复杂层次要求的策略.
  • 与最先进的方法相比,HPRS通过使用等级维护政策评估指标展示了更快的融合和更高的性能.
  • 除研究表明,HPRS在与安全和目标目标保持一致时有效地利用舒适性要求,并在冲突时无视它们.
  • 模拟现实实验表明,HPRS可以促进域名转移,而不需要手动的参数调整或调整.

结论:

  • HPRS提供了一种有效的自动化方法,用于合成符合复杂,层次要求的机器人政策.
  • 该方法提高了奖励信号的质量,从而提高了培训效率和政策绩效.
  • 通过自动平衡竞争目标,HPRS简化了设计过程,并显示了现实世界机器人技术的实际可行性.
  • 层次任务规范设计有助于机器人应用的强大的模拟到真实转移.