Jove
Visualize
联系我们
JoVE
x logofacebook logolinkedin logoyoutube logo
关于 JoVE
概览领导团队博客JoVE 帮助中心
作者
出版流程编辑委员会范围与政策同行评审常见问题投稿
图书馆员
用户评价订阅访问资源图书馆顾问委员会常见问题
研究
JoVE JournalMethods CollectionsJoVE Encyclopedia of Experiments存档
教育
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab Manual教师资源中心教师网站
使用条款与条件
隐私政策
政策

相关概念视频

Survival Tree01:19

Survival Tree

458
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
 Building a Survival Tree
Constructing a...
458
Decision Making: Traditional Method01:14

Decision Making: Traditional Method

5.6K
The process of hypothesis testing based on the traditional method includes calculating the critical value, testing the value of the test statistic using the sample data, and interpreting these values.
First, a specific claim about the population parameter is decided based on the research question and is stated in a simple form. Further, an opposing statement to this claim is also stated. These statements can act as null and alternative hypotheses, out of which a null hypothesis would be a...
5.6K
Cognitive Learning01:21

Cognitive Learning

1.5K
Cognitive learning is based on purposive behavior, incidental learning, and insight learning.
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
1.5K
Decision Making: P-value Method01:09

Decision Making: P-value Method

7.1K
The process of hypothesis testing based on the P-value method includes calculating the P- value using the sample data and interpreting it.
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim  is also stated. These statements can act as null and alternative hypotheses:  a null hypothesis would be a neutral statement while the alternative hypothesis can...
7.1K
Timing and Consequences on Behavior01:08

Timing and Consequences on Behavior

566
In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective. 
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
566
Law of Effect01:06

Law of Effect

5.1K
B.F. Skinner, a prominent figure in behavioral psychology, introduced operant conditioning by emphasizing the role of consequences in shaping behavior. This theory builds upon the law of effect proposed by Edward Thorndike, which posits that behaviors followed by satisfying outcomes are likely to be repeated. In contrast, those followed by unsatisfying outcomes are less likely to recur.
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...
5.1K

您也可能阅读

相关文章

通过共同作者、期刊和引用图与本文相关的文章。

排序
Same author

Wearable Data Economy: Implications of the FDA's 2026 General Wellness Policy.

Circulation. Population health and outcomes·2026
Same author

Qualitative Analysis of User Experiences of a mHealth Self-Care Intervention for Care Partners of Individuals with Traumatic Brain Injury.

Archives of rehabilitation research and clinical translation·2026
Same author

Is More Always Better With Digital Health Interventions? Shifting Engagement From Maximizing Use to Supporting Health.

Mayo Clinic proceedings. Digital health·2026
Same author

Reproducible workflow for online artificial intelligence in digital health.

Philosophical transactions. Series A, Mathematical, physical, and engineering sciences·2026
Same author

Classification of Plasmodium vivax as artefactual cells by the CellaVision DM9600: a system limitation in a non-endemic area.

Journal of microbiology, immunology, and infection = Wei mian yu gan ran za zhi·2026
Same author

scVIP: personalized modeling of single-cell transcriptomes for developmental and disease phenotypes.

bioRxiv : the preprint server for biology·2026

相关实验视频

Updated: Mar 13, 2026

Pavlovian Conditioned Approach Training in Rats
06:57

Pavlovian Conditioned Approach Training in Rats

Published on: February 4, 2016

11.6K

在强化学习中利用因果关系,用包装的决策时间来加强学习.

Daiqi Gao1, Hsin-Yu Lai2, Predrag Klasnja3

  • 1Harvard University.

Proceedings of machine learning research
|March 12, 2026
PubMed
概括

这项研究引入了一种新的在线强化学习 (RL) 方法,用于处理有袋决策时间的问题,有效地处理使用因果定向非循环图 (DAG) 的非马科维动态,以最大限度地提高移动健康应用中的奖励.

科学领域:

  • 人工智能的人工智能
  • 机器学习 机器学习
  • 计算科学 计算科学

背景情况:

  • 强化学习 (RL) 传统上假定马科维的决策过程.
  • 现实世界的问题,如移动健康干预,往往在决策时期 (袋子) 内表现出非马科夫式和非静止式的动态.
  • 现有的RL方法与这些复杂的时间依赖性作斗争.

研究的目的:

  • 开发一种在线强化学习 (RL) 算法,能够最大限度地提高决策时间和非马科夫转换的情况的奖励.
  • 为了应对共同优化一个袋子内的行动,共同影响一个奖励的挑战.
  • 适应RL用于周期性马尔科夫决策流程 (MDP),具有周期内非静止性.

主要方法:

  • 利用专家提供的因果定向非循环图 (DAG) 来建模袋子内的依赖关系.
  • 构建状态作为动态贝叶斯足够的历史数据的统计数据,以确保马科维的过渡.
  • 制定了一个周期性MDP的问题,并为在线RL优化推广了贝尔曼方程.
  • 评估了关于移动健康试验数据的拟议方法.

主要成果:

  • 拟议的状态构造方法确保了马科维亚状态在袋子内和跨袋的过渡.
  • 开发的在线RL算法有效地处理周期性MDPs中的非静态性.

更多相关视频

Operant Protocols for Assessing the Cost-benefit Analysis During Reinforced Decision Making by Rodents
07:05

Operant Protocols for Assessing the Cost-benefit Analysis During Reinforced Decision Making by Rodents

Published on: September 10, 2018

6.5K
An Automated T-maze Based Apparatus and Protocol for Analyzing Delay- and Effort-based Decision Making in Free Moving Rodents
07:42

An Automated T-maze Based Apparatus and Protocol for Analyzing Delay- and Effort-based Decision Making in Free Moving Rodents

Published on: August 2, 2018

14.5K

相关实验视频

Last Updated: Mar 13, 2026

Pavlovian Conditioned Approach Training in Rats
06:57

Pavlovian Conditioned Approach Training in Rats

Published on: February 4, 2016

11.6K
Operant Protocols for Assessing the Cost-benefit Analysis During Reinforced Decision Making by Rodents
07:05

Operant Protocols for Assessing the Cost-benefit Analysis During Reinforced Decision Making by Rodents

Published on: September 10, 2018

6.5K
An Automated T-maze Based Apparatus and Protocol for Analyzing Delay- and Effort-based Decision Making in Free Moving Rodents
07:42

An Automated T-maze Based Apparatus and Protocol for Analyzing Delay- and Effort-based Decision Making in Free Moving Rodents

Published on: August 2, 2018

14.5K
  • 构造状态被证明可以实现周期MDP的最大最佳值函数.
  • 在移动健康测试台上的成功评估证明了其实际适用性.
  • 结论:

    • 新的RL框架成功地解决了包装决策时间问题中的非马科夫式和非静态动态.
    • 基于DAG的状态构造提供了一种有效的方式来管理复杂的时间依赖.
    • 该方法为优化移动健康等领域的顺序决策提供了一个有希望的方法.