Jove
Visualize
联系我们
JoVE
x logofacebook logolinkedin logoyoutube logo
关于 JoVE
概览领导团队博客JoVE 帮助中心
作者
出版流程编辑委员会范围与政策同行评审常见问题投稿
图书馆员
用户评价订阅访问资源图书馆顾问委员会常见问题
研究
JoVE JournalMethods CollectionsJoVE Encyclopedia of Experiments存档
教育
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab Manual教师资源中心教师网站
使用条款与条件
隐私政策
政策

相关概念视频

Reinforcement Schedules01:24

Reinforcement Schedules

148
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
148
Timing and Consequences on Behavior01:08

Timing and Consequences on Behavior

97
In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective. 
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
97
Law of Effect01:06

Law of Effect

1.4K
B.F. Skinner, a prominent figure in behavioral psychology, introduced operant conditioning by emphasizing the role of consequences in shaping behavior. This theory builds upon the law of effect proposed by Edward Thorndike, which posits that behaviors followed by satisfying outcomes are likely to be repeated. In contrast, those followed by unsatisfying outcomes are less likely to recur.
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...
1.4K
Generalization, Discrimination, and Extinction01:24

Generalization, Discrimination, and Extinction

565
Generalization, discrimination, and extinction are key concepts in operant conditioning that influence how behaviors are learned and maintained.
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
565
BIBO stability of continuous and discrete -time systems01:24

BIBO stability of continuous and discrete -time systems

398
System stability is a fundamental concept in signal processing, often assessed using convolution. For a system to be considered bounded-input bounded-output (BIBO) stable, any bounded input signal must produce a bounded output signal. A bounded input signal is one where the modulus does not exceed a certain constant at any point in time.
To determine the BIBO stability, the convolution integral is utilized when a bounded continuous-time input is applied to a Linear Time-Invariant (LTI) system....
398
Decision Making01:20

Decision Making

112
Decision-making is a fundamental cognitive process that involves evaluating alternatives and selecting among them. This process can range from simple choices, such as deciding what to wear, to complex decisions, like choosing a major in college or a career path. The complexity of the decision often dictates the approach we use, which can be broadly categorized into two types: automatic and controlled decision-making.
Automatic decision-making is fast, intuitive, and relies on gut feelings...
112

您也可能阅读

相关文章

通过共同作者、期刊和引用图与本文相关的文章。

排序
Same author

Computer-Assisted Analytical Workflows for Natural Product Dereplication, Structure Elucidation, and Bioactivity-Oriented Prioritization.

Analytical chemistry·2026
Same author

Short-term nitric oxide gaseous treatment of post-rigor pork: Multidimensional impacts on meat quality, protein oxidation and proteolysis, and microstructure.

Food chemistry·2026
Same author

Awareness, Educational Needs, and Curriculum Preferences Regarding AI and Medical Big Data Education Among Clinical Medicine Undergraduates: Cross-Sectional Survey Study.

JMIR formative research·2026
Same author

Corrigendum to: Lithium Chloride Improves Electrophysiological and Memory Deficits in Rats with Streptozotocin-Induced Alzheimer's Disease.

Current Alzheimer research·2026
Same author

FedAttn-Credit: attention-augmented federated learning with adaptive differential privacy for rural inclusive finance credit assessment.

Scientific reports·2026
Same author

Author Correction: Multi-omics profiling reveals tumor microenvironment characteristics linked to immunotherapy response and prognosis in non-small cell lung cancer.

NPJ precision oncology·2026

相关实验视频

Updated: Jul 6, 2025

Operant Protocols for Assessing the Cost-benefit Analysis During Reinforced Decision Making by Rodents
07:05

Operant Protocols for Assessing the Cost-benefit Analysis During Reinforced Decision Making by Rodents

Published on: September 10, 2018

6.0K

半无限约束的马尔科夫决策过程和可证明有效的强化学习.

Liangyu Zhang, Yang Peng, Wenhao Yang

    IEEE transactions on pattern analysis and machine intelligence
    |January 1, 2024
    PubMed
    概括

    本研究介绍了半无限受约束的马尔科夫决策过程 (SICMDP) 和两个新的算法,SI-CMBRL和SI-CPO,用于使用强化学习解决复杂的控制任务.

    科学领域:

    • 人工智能的人工智能
    • 机器学习 机器学习
    • 运营研究 运营研究

    背景情况:

    • 约束马尔科夫决策流程 (CMDP) 被广泛用于在约束下进行顺序决策.
    • 现有的CMDP框架通常处理有限数量的约束,限制其适用性.
    • 需要一种能够解决决策问题的方法,这种方法具有连续性的约束.

    研究的目的:

    • 引入CMDPs的新型概括,称为半无限受约束的马尔科夫决策过程 (SICMDPs).
    • 开发和分析两个新的强化学习算法,SI-CMBRL和SI-CPO,为SICMDPs量身定制.
    • 为了证明这些算法的有效性在解决复杂的控制任务.

    主要方法:

    • 开发SI-CMBRL,一种基于模型的强化学习算法,将SICMDP转换为线性半无限编程 (LSIP) 问题.
    • 开发SI-CPO,一种政策优化算法,用于政策更新采用合作性随机近似.
    • 首次应用从半无限编程 (SIP) 到受约束强化学习的技术.

    主要成果:

    • 为SI-CMBRL和SI-CPO提供理论分析,包括代和样本复杂性.
    • 进行了广泛的数值实验,验证了SICMDP模型.
    • 证明了SI-CMBRL和SI-CPO在使用深度强化学习解决复杂的控制任务方面的能力.

    更多相关视频

    An Automated T-maze Based Apparatus and Protocol for Analyzing Delay- and Effort-based Decision Making in Free Moving Rodents
    07:42

    An Automated T-maze Based Apparatus and Protocol for Analyzing Delay- and Effort-based Decision Making in Free Moving Rodents

    Published on: August 2, 2018

    13.6K
    Three Laboratory Procedures for Assessing Different Manifestations of Impulsivity in Rats
    09:12

    Three Laboratory Procedures for Assessing Different Manifestations of Impulsivity in Rats

    Published on: March 17, 2019

    9.5K

    相关实验视频

    Last Updated: Jul 6, 2025

    Operant Protocols for Assessing the Cost-benefit Analysis During Reinforced Decision Making by Rodents
    07:05

    Operant Protocols for Assessing the Cost-benefit Analysis During Reinforced Decision Making by Rodents

    Published on: September 10, 2018

    6.0K
    An Automated T-maze Based Apparatus and Protocol for Analyzing Delay- and Effort-based Decision Making in Free Moving Rodents
    07:42

    An Automated T-maze Based Apparatus and Protocol for Analyzing Delay- and Effort-based Decision Making in Free Moving Rodents

    Published on: August 2, 2018

    13.6K
    Three Laboratory Procedures for Assessing Different Manifestations of Impulsivity in Rats
    09:12

    Three Laboratory Procedures for Assessing Different Manifestations of Impulsivity in Rats

    Published on: March 17, 2019

    9.5K

    结论:

    • SICMDPs为决策提供了一个强大的框架,具有持续的约束.
    • 建议的SI-CMBRL和SI-CPO算法对于解决SICMDPs是有效的,理论上是合理的.
    • 这项工作开创了半无限编程在受约束强化学习中的应用.