Related Experiment Video
Updated: Feb 24, 2026

Measuring the Subjective Value of Risky and Ambiguous Options using Experimental Economics and Functional MRI Methods
Published on: September 19, 2012
Optimism in the face of uncertainty supported by a statistically-designed multi-armed bandit algorithm
1Graduate School of Science and Engineering, Tokyo Denki University, Japan; School of Science and Engineering, Tokyo Denki University, Japan.
The unified Overtaking method enhances multi-armed bandit problem-solving by redefining value functions. This new approach, based on statistical confidence intervals, outperforms the UCB algorithm for exponentially distributed rewards.
Area of Science:
- Decision Sciences
- Machine Learning
- Optimization Theory
Background:
- Sequential decision-making problems often involve uncertainty.
- The Overtaking method is an effective algorithm for multi-armed bandit problems, previously defined by heuristic patterns.
- Optimism in the face of uncertainty is a key principle in these problems.
Purpose of the Study:
- To redefine and unify the value functions of the Overtaking method.
- To enhance the universality and applicability of the Overtaking method.
- To propose a statistics-based interpretation of optimism in sequential decision-making.
Main Methods:
- Unified formulation of the Overtaking method using upper confidence bounds on expected rewards.
- Development of the Overtaking method for exponentially distributed rewards.
- Numerical analysis and comparison with the UCB algorithm.
Main Results:
- The unified Overtaking method demonstrates enhanced universality.
- A novel Overtaking method for exponentially distributed rewards was developed and analyzed.
- The proposed method statistically outperforms the UCB algorithm on average.
Conclusions:
- The principle of optimism in sequential decision-making can be viewed as a statistical consequence of the law of large numbers.
- Unified formulations improve algorithm universality in multi-armed bandit problems.
- The new Overtaking method offers improved performance for specific reward distributions.
Related Concept Videos
Unrealistic Optimism Bias
Uncertainty: Confidence Intervals
Propagation of Uncertainty from Random Error
Uncertainty: Overview
Propagation of Uncertainty from Systematic Error
Randomized Experiments
Simple randomization
Simple...

