Jove
Visualize
联系我们
JoVE
x logofacebook logolinkedin logoyoutube logo
关于 JoVE
概览领导团队博客JoVE 帮助中心
作者
出版流程编辑委员会范围与政策同行评审常见问题投稿
图书馆员
用户评价订阅访问资源图书馆顾问委员会常见问题
研究
JoVE JournalMethods CollectionsJoVE Encyclopedia of Experiments存档
教育
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab Manual教师资源中心教师网站
使用条款与条件
隐私政策
政策

相关概念视频

Observational Learning01:12

Observational Learning

791
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
791
Associative Learning01:27

Associative Learning

1.2K
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
1.2K
Reinforcement Schedules01:24

Reinforcement Schedules

436
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
436
Generalization, Discrimination, and Extinction01:24

Generalization, Discrimination, and Extinction

1.3K
Generalization, discrimination, and extinction are key concepts in operant conditioning that influence how behaviors are learned and maintained.
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
1.3K
Reinforcement01:23

Reinforcement

786
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
786
Nonconscious Mimicry01:13

Nonconscious Mimicry

5.1K
Nonconscious mimicry occurs when individuals alter their mannerisms to match the behaviors and expressions of those nearby, without intention.
5.1K

您也可能阅读

相关文章

通过共同作者、期刊和引用图与本文相关的文章。

排序
Same author

Kinetic Analysis and Starch Digestion Product Composition Reveal the Subtle Relationship between the Anthocyanidin Structure and Inhibitory Activity on Pancreatic α-Amylase.

Journal of agricultural and food chemistry·2025
Same author

Performance of electro-assisted ecological floating bed in antibiotics and conventional pollutants degradation: Mechanisms and microbial response.

Journal of environmental management·2025
Same author

Flexible OLED Performance Enhancement: The Impact of Ag NWs: Ag NPs Electrode-Integrated MoO<sub></sub> QDs Hole-Injection Layer.

ACS applied materials & interfaces·2025
Same author

Effects of Replacing Inorganic Sources of Copper, Manganese, and Zinc with Different Organic Forms on Mineral Status, Immune Biomarkers, and Lameness of Lactating Cows.

Animals : an open access journal from MDPI·2025
Same author

Dynamic single cell transcriptomics defines kidney FGF23/KL bioactivity and novel segment-specific inflammatory targets.

Kidney international·2025
Same author

Single cell approaches define neural stem cell niches and identify microglial ligands that can enhance precursor-mediated oligodendrogenesis.

Cell reports·2025

相关实验视频

Updated: Jan 8, 2026

A Step-by-Step Implementation of DeepBehavior, Deep Learning Toolbox for Automated Behavior Analysis
05:41

A Step-by-Step Implementation of DeepBehavior, Deep Learning Toolbox for Automated Behavior Analysis

Published on: February 6, 2020

9.8K

通过基于动态的行为相似性的学习表征,用于深度强化学习.

Dayang Liang1, Yunlong Liu1

  • 1Department of Automation, Xiamen University, Xiamen, 361005, China.

Neural networks : the official journal of the International Neural Network Society
|December 21, 2025
PubMed
概括

我们介绍了基于动态的行为相似性 (RDS) 的表示学习,以改善深度强化学习. 通过消除奖励依赖,RDS增强了表示学习,在复杂的操纵任务中实现了显著的性能增长.

科学领域:

  • 人工智能的人工智能
  • 机器学习 机器学习
  • 机器人技术 机器人技术 机器人技术

背景情况:

  • 深度强化学习需要从视觉数据中学习与任务相关的表示.
  • 行为相似度指标对等状态进行组合,但由于奖励稀少,代表性崩.
  • 这限制了复杂应用程序的可扩展性.

研究的目的:

  • 提出一种新的方法,即基于动态的行为相似性 (RDS) 的表示学习,以克服现有的表示学习方法的局限性.
  • 开发一种独立于奖励的相似度指标,以保持行为歧视性,以改善深度强化学习.

主要方法:

  • 引入了一个动态驱动的相似度指标,消除了奖励依赖.
  • 集成的动态过渡距离与可训练的高斯噪声来缓解度量降解.
  • 利用潜伏轨迹距离来量化任务差异和提取相关特征.

主要成果:

  • 在复杂的DeepMind Control,MetaWorld和Aroit操纵任务上,RDS表现出比基线方法更好的性能.
  • 与DrQ-v2和最先进的方法相比,分别实现了43%和30%的显著改进.
  • 废弃性研究证实了RDS方法中的单个成分的有效性.
关键词:
行为相似度指标.深度强化学习的学习.很少有奖励奖励.与任务相关的表示.

更多相关视频

The Double-H Maze: A Robust Behavioral Test for Learning and Memory in Rodents
09:01

The Double-H Maze: A Robust Behavioral Test for Learning and Memory in Rodents

Published on: July 8, 2015

13.1K
Decoding Natural Behavior from Neuroethological Embedding
08:00

Decoding Natural Behavior from Neuroethological Embedding

Published on: October 3, 2025

556

相关实验视频

Last Updated: Jan 8, 2026

A Step-by-Step Implementation of DeepBehavior, Deep Learning Toolbox for Automated Behavior Analysis
05:41

A Step-by-Step Implementation of DeepBehavior, Deep Learning Toolbox for Automated Behavior Analysis

Published on: February 6, 2020

9.8K
The Double-H Maze: A Robust Behavioral Test for Learning and Memory in Rodents
09:01

The Double-H Maze: A Robust Behavioral Test for Learning and Memory in Rodents

Published on: July 8, 2015

13.1K
Decoding Natural Behavior from Neuroethological Embedding
08:00

Decoding Natural Behavior from Neuroethological Embedding

Published on: October 3, 2025

556

结论:

  • 使用基于动态的行为相似性 (RDS) 的表示学习有效地解决了深度强化学习中的表示崩.
  • 拟议的方法通过利用基于动态的相似性来增强对任务相关特征的学习,在具有挑战性的机器人任务上表现出强的表现.