Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Observational Learning01:12

Observational Learning

142
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
142
Associative Learning01:27

Associative Learning

300
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
300
Purposive Learning01:22

Purposive Learning

101
E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a...
101
Band Theory02:35

Band Theory

15.0K
When two or more atoms come together to form a molecule, their atomic orbitals combine and molecular orbitals of distinct energies result. In a solid, there are a large number of atoms, and therefore a large number of atomic orbitals that may be combined into molecular orbitals. These groups of molecular orbitals are so closely placed together to form continuous regions of energies, known as the bands.
The energy difference between these bands is known as the band gap.
Conductor, Semiconductor,...
15.0K
Cognitive Learning01:21

Cognitive Learning

222
Cognitive learning is based on purposive behavior, incidental learning, and insight learning.
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
222
Bandpass Sampling01:17

Bandpass Sampling

164
In signal processing, bandpass sampling is an effective technique for sampling signals that have most of their energy concentrated within a narrow frequency band. This type of signal is known as a bandpass signal. The key principle of bandpass sampling involves sampling the signal at a rate that is greater than twice the signal's bandwidth to prevent aliasing.
A bandpass signal has a spectrum with a lower frequency limit, denoted as ω1, and an upper frequency limit, denoted as ω2....
164

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Mobile intervention for emerging adults with regular cannabis use: a micro-randomized trial.

Lancet regional health. Americas·2026
Same author

Reproducible workflow for online artificial intelligence in digital health.

Philosophical transactions. Series A, Mathematical, physical, and engineering sciences·2026
Same author

A Beginner's Guide to Applying Large Language Models in Behavioral Interventions.

JMIR mHealth and uHealth·2026
Same author

Evaluation of the construct validity of the Michigan Fatigability Index (MIFI) short forms: a cross-sectional survey study.

Journal of patient-reported outcomes·2026
Same author

Non-Stationary Latent Auto-Regressive Bandits.

Reinforcement learning journal·2026
Same author

When and Why Hyperbolic Discounting Matters for Reinforcement Learning Interventions.

Reinforcement learning journal·2026

Related Experiment Video

Updated: Jun 9, 2025

Using MazeSuite and Functional Near Infrared Spectroscopy to Study Learning in Spatial Navigation
20:12

Using MazeSuite and Functional Near Infrared Spectroscopy to Study Learning in Spatial Navigation

Published on: October 8, 2011

30.5K

Online learning in bandits with predicted context.

Yongyi Guo1, Ziping Xu2, Susan Murphy2

  • 1Department of Statistics, University of Wisconsin-Madison, Madison, WI, 53706.

Proceedings of Machine Learning Research
|October 28, 2024
PubMed
Summary

This study introduces a novel online algorithm for contextual bandit problems with noisy context data. The new approach achieves sublinear regret, outperforming classical algorithms in scenarios with unobserved true contexts.

More Related Videos

Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
08:05

Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques

Published on: June 30, 2020

7.5K
Project-Based Learning Guidelines for Health Sciences Students: An Analysis with Data Mining and Qualitative Techniques
13:44

Project-Based Learning Guidelines for Health Sciences Students: An Analysis with Data Mining and Qualitative Techniques

Published on: December 9, 2022

3.5K

Related Experiment Videos

Last Updated: Jun 9, 2025

Using MazeSuite and Functional Near Infrared Spectroscopy to Study Learning in Spatial Navigation
20:12

Using MazeSuite and Functional Near Infrared Spectroscopy to Study Learning in Spatial Navigation

Published on: October 8, 2011

30.5K
Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
08:05

Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques

Published on: June 30, 2020

7.5K
Project-Based Learning Guidelines for Health Sciences Students: An Analysis with Data Mining and Qualitative Techniques
13:44

Project-Based Learning Guidelines for Health Sciences Students: An Analysis with Data Mining and Qualitative Techniques

Published on: December 9, 2022

3.5K

Area of Science:

  • Machine Learning
  • Reinforcement Learning
  • Statistical Decision Theory

Background:

  • The contextual bandit problem is crucial for sequential decision-making under uncertainty.
  • Classical bandit algorithms often assume perfect knowledge of the context, which is unrealistic in many applications.
  • Existing methods fail when context information is noisy or predicted by another model, leading to non-vanishing errors.

Purpose of the Study:

  • To develop the first online algorithm capable of achieving sublinear regret in contextual bandit problems with noisy context data.
  • To address settings where only a predicted, rather than true, context is available for decision-making.
  • To provide theoretical guarantees for the proposed algorithm under mild conditions.

Main Methods:

  • Extension of the classical statistical measurement error model to the online decision-making framework.
  • Development of a novel online algorithm designed to handle policies dependent on noisy context observations.
  • Theoretical analysis to establish sublinear regret guarantees.

Main Results:

  • The proposed algorithm achieves sublinear regret, a significant improvement over classical methods in noisy context settings.
  • The algorithm demonstrates effectiveness even when context errors are substantial and do not diminish over time.
  • Validation through simulations using both synthetic and real-world digital intervention datasets.

Conclusions:

  • The novel algorithm provides a robust solution for contextual bandit problems with unobserved or predicted contexts.
  • This work advances the field by offering the first theoretical guarantees for sublinear regret in this challenging setting.
  • The approach has practical implications for various applications relying on machine learning predictions for decision-making.