Related Experiment Video
Updated: Jun 19, 2025

Measuring the Subjective Value of Risky and Ambiguous Options using Experimental Economics and Functional MRI Methods
Published on: September 19, 2012
Thompson Sampling for Stochastic Bandits with Noisy Contexts: An Information-Theoretic Regret Analysis
Sharu Theresa Jose1, Shana Moothedath2
1School of Computer Science, University of Birmingham, Birmingham B15 2TT, UK.
This study introduces a modified Thompson sampling algorithm for noisy contextual bandits (CB). The algorithm approximates an oracle policy, achieving near-optimal Bayesian cumulative regret scaling as O˜(mT) for Gaussian bandits.
Area of Science:
- Machine Learning
- Reinforcement Learning
- Information Theory
Background:
- Contextual bandits (CB) are crucial for sequential decision-making with partial feedback.
- Real-world CB often involve noisy context observations, complicating policy design.
- Existing methods struggle with unknown noise channel parameters in CB.
Purpose of the Study:
- To design an effective action policy for stochastic linear contextual bandits with noisy context observations.
- To approximate the performance of a Bayesian oracle with access to the true context and reward model.
- To analyze the Bayesian cumulative regret of the proposed policy using information-theoretic tools.
Main Methods:
- Introduced a modified Thompson sampling algorithm tailored for noisy CB.
- Employed information-theoretic analysis to derive regret bounds.
- Investigated the impact of delayed context information on regret.
- Conducted empirical evaluations against established baseline algorithms.
Main Results:
- The proposed algorithm achieves Bayesian cumulative regret scaling as O˜(mT) for Gaussian bandits with Gaussian context noise under specific prior variance conditions.
- Demonstrated that delayed true context observations can lead to reduced regret.
- Empirical results validate the algorithm's performance against baselines.
Conclusions:
- The modified Thompson sampling algorithm offers a robust solution for contextual bandits with noisy contexts.
- The findings provide theoretical guarantees and practical insights into managing context uncertainty in CB.
- Delayed context information presents a potential avenue for improving CB performance.
Related Concept Videos
Propagation of Uncertainty from Random Error
Propagation of Uncertainty from Systematic Error
Sampling Theorem
Noncompartmental Analysis: Statistical Moment Theory
Sampling Distribution
Randomized Experiments
Simple randomization
Simple...

