Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Decision Making: P-value Method01:09

Decision Making: P-value Method

7.3K
The process of hypothesis testing based on the P-value method includes calculating the P- value using the sample data and interpreting it.
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim  is also stated. These statements can act as null and alternative hypotheses:  a null hypothesis would be a neutral statement while the alternative hypothesis can...
7.3K
Expected Value01:15

Expected Value

8.3K
The expected value is known as the "long-term" average or mean. This means that over the long term of experimenting over and over, you would expect this average. The expected average is represented by the symbol μ. It is calculated as follows:
8.3K
Multi-input and Multi-variable systems01:22

Multi-input and Multi-variable systems

504
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence of...
504
Optimal Foraging00:48

Optimal Foraging

14.2K
How animals obtain and eat their food is called foraging behavior. Foraging can include searching for plants and hunting for prey and depends on the species and environment.
14.2K
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving01:29

Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving

407
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
407
Bias01:22

Bias

8.0K
Bias refers to any tendency that prevents a question from being considered unprejudiced. In research, bias occurs when one outcome or answer is selected or encouraged over others in sampling or testing. Bias can occur during any research phase, including study design, data collection, analysis, and publication.
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
8.0K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Optimism in the face of uncertainty supported by a statistically-designed multi-armed bandit algorithm.

Bio Systems·2017
See all related articles

Related Experiment Video

Updated: Apr 7, 2026

Measuring the Subjective Value of Risky and Ambiguous Options using Experimental Economics and Functional MRI Methods
13:04

Measuring the Subjective Value of Risky and Ambiguous Options using Experimental Economics and Functional MRI Methods

Published on: September 19, 2012

12.5K

Overtaking method based on sand-sifter mechanism: Why do optimistic value functions find optimal solutions in

Kento Ochi1, Moto Kamiura2

  • 1Graduate School of Science and Engineering, Tokyo Denki University, Japan.

Bio Systems
|July 14, 2015
PubMed
Summary

A new Overtaking method for multi-armed bandit problems uses an optimistic value function. This approach achieves high accuracy and low regret, potentially outperforming existing UCB algorithms in uncertain environments.

Keywords:
Confidence intervalExploration–exploitation dilemmaMulti-armed bandit problemOptimismUCB algorithm

More Related Videos

A Tactile Automated Passive-Finger Stimulator TAPS
19:44

A Tactile Automated Passive-Finger Stimulator TAPS

Published on: June 3, 2009

14.3K
The Joint Effect of Social Comparison and Social Distance on Evaluation of Intertemporal Choice Outcomes in Event-related Potential Studies
08:24

The Joint Effect of Social Comparison and Social Distance on Evaluation of Intertemporal Choice Outcomes in Event-related Potential Studies

Published on: August 25, 2023

1.3K

Related Experiment Videos

Last Updated: Apr 7, 2026

Measuring the Subjective Value of Risky and Ambiguous Options using Experimental Economics and Functional MRI Methods
13:04

Measuring the Subjective Value of Risky and Ambiguous Options using Experimental Economics and Functional MRI Methods

Published on: September 19, 2012

12.5K
A Tactile Automated Passive-Finger Stimulator TAPS
19:44

A Tactile Automated Passive-Finger Stimulator TAPS

Published on: June 3, 2009

14.3K
The Joint Effect of Social Comparison and Social Distance on Evaluation of Intertemporal Choice Outcomes in Event-related Potential Studies
08:24

The Joint Effect of Social Comparison and Social Distance on Evaluation of Intertemporal Choice Outcomes in Event-related Potential Studies

Published on: August 25, 2023

1.3K

Area of Science:

  • Machine Learning
  • Reinforcement Learning
  • Decision Theory

Background:

  • Multi-armed bandit problems involve agents selecting optimal actions from multiple options with random rewards.
  • The Upper Confidence Bound (UCB) algorithm balances exploration and exploitation for logarithmic regret.
  • Empirical evidence suggests optimistic value functions perform well, though the underlying reasons are not fully understood.

Purpose of the Study:

  • To propose a novel method, the Overtaking method, for solving multi-armed bandit problems.
  • To investigate the performance benefits of optimism in agent decision-making within uncertain environments.
  • To introduce a new value function definition based on confidence interval upper bounds.

Main Methods:

  • The Overtaking method defines a value function as an upper bound of the confidence interval for the expected reward.
  • This value function asymptotically approaches the true expected reward from above.
  • A 'sand-sifter mechanism' prevents suboptimal arms from regaining value, ensuring focus on the current best arm.

Main Results:

  • The Overtaking method demonstrates high accuracy and low regret in multi-armed bandit tasks.
  • Certain configurations of the Overtaking method's value functions outperform traditional UCB algorithms.
  • The method ensures the agent is almost certain to identify the optimal arm under asymptotic conditions.

Conclusions:

  • Optimism in agent value functions provides significant advantages in uncertain environments, as exemplified by the Overtaking method.
  • The proposed Overtaking method offers a competitive alternative to UCB algorithms, particularly in its ability to converge on the optimal arm.
  • This research highlights the efficacy of confidence-interval-based optimism for efficient decision-making in bandit problems.