Related Experiment Video
Updated: Aug 2, 2025

WheelCon: A Wheel Control-Based Gaming Platform for Studying Human Sensorimotor Control
Published on: August 15, 2020
Quantifying the stochasticity of policy parameters in reinforcement learning problems
Vahe Galstyan1,2, David B Saakian2
1AMOLF, Science Park 104, 1098 XG Amsterdam, Netherlands.
Abstract:
The stochastic dynamics of reinforcement learning is studied using a master equation formalism. We consider two different problems-Q learning for a two-agent game and the multiarmed bandit problem with policy gradient as the learning method. The master equation is constructed by introducing a probability distribution over continuous policy parameters or over both continuous policy parameters and discrete state variables (a more advanced case). We use a version of the moment closure approximation to solve for the stochastic dynamics of the models. Our method gives accurate estimates for the mean and the (co)variance of policy variables. For the case of the two-agent game, we find that the variance terms are finite at steady state and derive a system of algebraic equations for computing them directly.
More Related Videos
Related Concept Videos
Propagation of Uncertainty from Random Error
Random Variables
Uppercase letters such as X or Y denote a random variable. Lowercase letters like x or y denote the value of a random variable. If X is a random variable, then X is written in words, and x is given as a number.
For example, let X = the...
Reinforcement Schedules
Once a behavior is learned,...
Propagation of Uncertainty from Systematic Error
Randomized Experiments
Simple randomization
Simple...
Parametric Survival Analysis: Weibull and Exponential Methods
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...

