Jove
Visualize
联系我们
JoVE
x logofacebook logolinkedin logoyoutube logo
关于 JoVE
概览领导团队博客JoVE 帮助中心
作者
出版流程编辑委员会范围与政策同行评审常见问题投稿
图书馆员
用户评价订阅访问资源图书馆顾问委员会常见问题
研究
JoVE JournalMethods CollectionsJoVE Encyclopedia of Experiments存档
教育
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab Manual教师资源中心教师网站
使用条款与条件
隐私政策
政策

相关概念视频

Observational Learning01:12

Observational Learning

841
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
841
Purposive Learning01:22

Purposive Learning

447
E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a...
447
Associative Learning01:27

Associative Learning

1.3K
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
1.3K
Avoidance Learning and Learned Helplessness01:14

Avoidance Learning and Learned Helplessness

2.5K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
2.5K
Optimal Foraging00:48

Optimal Foraging

13.6K
How animals obtain and eat their food is called foraging behavior. Foraging can include searching for plants and hunting for prey and depends on the species and environment.
13.6K
Reinforcement01:23

Reinforcement

841
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
841

您也可能阅读

相关文章

通过共同作者、期刊和引用图与本文相关的文章。

排序
Same author

M phase phosphorylation of the epigenetic regulator UHRF1 regulates its physical association with the deubiquitylase USP7 and stability.

Proceedings of the National Academy of Sciences of the United States of America·2012
Same author

Biosynthesis of ethyl oleate, a primer pheromone, in the honey bee (Apis mellifera L.).

Insect biochemistry and molecular biology·2012
Same author

Co-delivery strategies based on multifunctional nanocarriers for cancer therapy.

Current drug metabolism·2012
Same author

Efficacy of gemifloxacin for the treatment of experimental Staphylococcus aureus keratitis.

Journal of ocular pharmacology and therapeutics : the official journal of the Association for Ocular Pharmacology and Therapeutics·2012
Same author

Expression profile analysis of the polygalacturonase-inhibiting protein genes in rice and their responses to phytohormones and fungal infection.

Plant cell reports·2012
Same author

Characterizing natural dissolved organic matter in a freshly submerged catchment (Three Gorges Dam, China) using UV absorption, fluorescence spectroscopy and PARAFAC.

Water science and technology : a journal of the International Association on Water Pollution Research·2012

相关实验视频

Updated: Jan 17, 2026

Investigating Motor Skill Learning Processes with a Robotic Manipulandum
07:52

Investigating Motor Skill Learning Processes with a Robotic Manipulandum

Published on: February 12, 2017

9.2K

在多代理强化学习中为授权合作行动量身定制知识.

Hu Fu1, Yihua Tan1, Hao Chen2

  • 1School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Luoyu Road 1037, Wuhan, 430074, Hubei, China.

Neural networks : the official journal of the International Neural Network Society
|September 19, 2025
PubMed
概括

为授权合作行动定制知识 (TKCA) 增强了多代理强化学习 (MARL),使代理人能够选择特定的知识,克服部分参数共享的局限性,以实现更好的协作.

关键词:
行为多样性 行为多样性知识编码器知识编码器知识选择器知识选择器多个代理强化学习学习多个代理强化学习学习

更多相关视频

Real-Time Proxy-Control of Re-Parameterized Peripheral Signals using a Close-Loop Interface
11:54

Real-Time Proxy-Control of Re-Parameterized Peripheral Signals using a Close-Loop Interface

Published on: May 8, 2021

5.1K
The HoneyComb Paradigm for Research on Collective Human Behavior
06:48

The HoneyComb Paradigm for Research on Collective Human Behavior

Published on: January 19, 2019

9.8K

相关实验视频

Last Updated: Jan 17, 2026

Investigating Motor Skill Learning Processes with a Robotic Manipulandum
07:52

Investigating Motor Skill Learning Processes with a Robotic Manipulandum

Published on: February 12, 2017

9.2K
Real-Time Proxy-Control of Re-Parameterized Peripheral Signals using a Close-Loop Interface
11:54

Real-Time Proxy-Control of Re-Parameterized Peripheral Signals using a Close-Loop Interface

Published on: May 8, 2021

5.1K
The HoneyComb Paradigm for Research on Collective Human Behavior
06:48

The HoneyComb Paradigm for Research on Collective Human Behavior

Published on: January 19, 2019

9.8K

科学领域:

  • 人工智能的人工智能
  • 机器学习 机器学习
  • 多代理系统 多代理系统

背景情况:

  • 在多代理强化学习 (MARL) 中,有效的协作依赖于行为多样性.
  • 目前的方法使用部分参数共享,造成培训冲突和知识冗余,因为不同的代理需求.
  • 这种方法限制了复杂的MARL场景中的可扩展性和性能.

研究的目的:

  • 引入一种新的方法,即为授权合作行动 (TKCA) 量身定制知识,以解决 MARL. 的局限性.
  • 为了使代理人能够获得和利用环境特定的知识,以改善决策.
  • 在合作MARL中平衡行为多样性与算法可扩展性.

主要方法:

  • TKCA使用知识编码器来处理各种环境知识.
  • 知识选择器网络帮助个人代理人选择相关的知识进行决策.
  • 这种方法为每个代理方便定制的知识获取.

主要成果:

  • 与现有的方法相比,TKCA在挑战StarCraftII微管理游戏方面表现出更好的表现.
  • 谷歌研究足球比赛中的评估也证实了TKCA方法的有效性.
  • 提出的方法成功地改善了复杂的MARL环境中的协作策略.

结论:

  • 在MARL中,TKCA有效地解决了培训冲突和知识冗余问题,通过将知识定制为个人代理.
  • 这种方法增强了代理人的决策能力,并促进了行为多样性,以改善合作.
  • TKCA为推进可扩展和高性能MARL系统提供了一个有前途的方向.