Jove
Visualize
联系我们
JoVE
x logofacebook logolinkedin logoyoutube logo
关于 JoVE
概览领导团队博客JoVE 帮助中心
作者
出版流程编辑委员会范围与政策同行评审常见问题投稿
图书馆员
用户评价订阅访问资源图书馆顾问委员会常见问题
研究
JoVE JournalMethods CollectionsJoVE Encyclopedia of Experiments存档
教育
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab Manual教师资源中心教师网站
使用条款与条件
隐私政策
政策

相关概念视频

Optimal Foraging00:48

Optimal Foraging

13.6K
How animals obtain and eat their food is called foraging behavior. Foraging can include searching for plants and hunting for prey and depends on the species and environment.
13.6K
Observational Learning01:12

Observational Learning

841
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
841
Reinforcement01:23

Reinforcement

841
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
841
Randomized Experiments01:13

Randomized Experiments

8.9K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
8.9K
Reinforcement Schedules01:24

Reinforcement Schedules

460
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
460
Decision Making: P-value Method01:09

Decision Making: P-value Method

6.8K
The process of hypothesis testing based on the P-value method includes calculating the P- value using the sample data and interpreting it.
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim  is also stated. These statements can act as null and alternative hypotheses:  a null hypothesis would be a neutral statement while the alternative hypothesis can...
6.8K

您也可能阅读

相关文章

通过共同作者、期刊和引用图与本文相关的文章。

排序
Same author

Biomarker screen-guided care for preterm birth risk in nulliparous pregnancies: a subgroup analysis of the PRIME randomized controlled trial.

The journal of maternal-fetal & neonatal medicine : the official journal of the European Association of Perinatal Medicine, the Federation of Asia and Oceania Perinatal Societies, the International Society of Perinatal Obstetricians·2026
Same author

Scaling Up Bayesian Neural Networks with Neural Networks.

Transactions on machine learning research·2026
Same author

Longitudinal Analysis of Peripheral MicroRNA Expression and Depressive Symptom Severity Change in a Community Cohort.

Epigenomes·2026
Same author

Statistics and AI - A Fireside Conversation.

Harvard data science review·2026
Same author

A Bayesian Time-Varying Psychophysiological Interaction Model.

Data science in science·2026
Same author

Neurodatascience: Past, Present, and Future.

Data science in science·2026

相关实验视频

Updated: Jan 17, 2026

Operant Protocols for Assessing the Cost-benefit Analysis During Reinforced Decision Making by Rodents
07:05

Operant Protocols for Assessing the Cost-benefit Analysis During Reinforced Decision Making by Rodents

Published on: September 10, 2018

6.4K

强化个人最佳政策的学习,从异构的数据.

By Rui Miao1, Babak Shahbaba2, Annie Qu2

  • 1National Heart, Lung, and Blood Institute.

Annals of statistics
|September 18, 2025
PubMed
概括

这项研究引入了使用异质数据进行个性化离线强化学习 (RL) 的新框架. 被惩罚的悲观的个性化政策学习 (P4L) 算法优化了针对不同人群的政策,优于现有的方法.

科学领域:

  • 人工智能的人工智能
  • 机器学习 机器学习
  • 强化学习是一种强化学习.

背景情况:

  • 线下强化学习 (RL) 通过预先收集的数据寻找最佳的政策.
  • 从异质数据中学习是线下RL的一个关键挑战.
  • 现有的方法往往会给异质人群带来低于最佳的政策.

研究的目的:

  • 为异质的时间静止马尔科夫决策过程 (MDP) 提出个性化的离线政策优化框架.
  • 解决处理不同人口数据的传统方法的局限性.
  • 开发一个算法,有效地估计个别的Q函数.

主要方法:

  • 开发了一个具有单个潜在变量的异质模型.
  • 引入了惩罚性的悲观的个性化政策学习 (P4L) 算法.
  • 在部分覆盖假设下,平均遗憾率保证快速.

主要成果:

  • 拟议的框架有效地估计了个别的Q函数.
  • P4L算法在模拟和现实应用中展示了卓越的数值性能.
  • 与现有方法相比,在异质人群中实现了更好的政策优化.

结论:

关键词:
90C4040 没有任何问题.91B69 它们是什么?动态处理方案 动态处理方案不同质的数据 异质的数据马尔科夫决策过程精确学习是指精确的学习.初级 62C2020 在线 62C20

相关实验视频

Last Updated: Jan 17, 2026

Operant Protocols for Assessing the Cost-benefit Analysis During Reinforced Decision Making by Rodents
07:05

Operant Protocols for Assessing the Cost-benefit Analysis During Reinforced Decision Making by Rodents

Published on: September 10, 2018

6.4K
  • 个性化的离线政策优化框架有效地处理MDP中的异质数据.
  • P4L算法为个性化政策学习提供了一个强大的解决方案.
  • 该研究强调了在线RL对不同人群的个性化方法的重要性.