Related Experiment Video
Updated: Nov 27, 2025

08:05
Design and Analysis for Fall Detection System Simplification
Published on: April 6, 2020
11.0K
On Gap-Based Lower Bounding Techniques for Best-Arm Identification.
Lan V Truong1, Jonathan Scarlett2
1Department of Engineering, University of Cambridge, Cambridge CB2 1PZ, UK.
Entropy (Basel, Switzerland)
|December 8, 2020
Summary
This study refines gap-based methods for multi-armed bandit problems, showing they match newer divergence-based techniques for best-arm identification lower bounds. Both approaches effectively determine sample complexity for identifying the optimal arm.
Area of Science:
- Machine Learning
- Reinforcement Learning
- Optimization
Background:
- The multi-armed bandit problem is a fundamental challenge in sequential decision-making.
- Establishing lower bounds on sample complexity is crucial for efficient algorithm design.
- Previous work introduced divergence-based and gap-based approaches for best-arm identification.
Purpose of the Study:
- To analyze and refine techniques for establishing lower bounds on arm pulls for best-arm identification.
- To compare the effectiveness of divergence-based and gap-based approaches.
- To investigate the performance of these methods under Bernoulli rewards.
Main Methods:
- Theoretical analysis of existing divergence-based and gap-based methods.
- Refinement of the gap-based approach for Bernoulli rewards.
- Comparison of lower bounds derived from both techniques.
Main Results:
- The refined gap-based approach matches the divergence-based approach (up to constant factors) for Bernoulli rewards.
- This holds true even when rewards are bounded away from zero and one.
- Both methods are shown to be effective for determining sample complexity lower bounds.
Conclusions:
- Refined gap-based methods offer comparable performance to divergence-based methods for best-arm identification.
- Both approaches provide valuable insights into the sample complexity of identifying the best arm.
- These findings contribute to a better understanding of optimal strategies in multi-armed bandit settings.

