Related Experiment Video
Updated: Mar 13, 2026

Pavlovian Conditioned Approach Training in Rats
Published on: February 4, 2016
Harnessing Causality in Reinforcement Learning With Bagged Decision Times.
Daiqi Gao1, Hsin-Yu Lai2, Predrag Klasnja3
1Harvard University.
This study introduces a novel online reinforcement learning (RL) approach for problems with bagged decision times, effectively handling non-Markovian dynamics using causal directed acyclic graphs (DAGs) to maximize rewards in mobile health applications.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Computational Science
Background:
- Reinforcement learning (RL) traditionally assumes Markovian decision processes.
- Real-world problems, such as mobile health interventions, often exhibit non-Markovian and non-stationary dynamics within decision periods (bags).
- Existing RL methods struggle with these complex temporal dependencies.
Purpose of the Study:
- To develop an online reinforcement learning (RL) algorithm capable of maximizing rewards in scenarios with bagged decision times and non-Markovian transitions.
- To address the challenge of jointly optimizing actions within a bag that collectively influence a single reward.
- To adapt RL for periodic Markov decision processes (MDPs) with intra-period non-stationarity.
Main Methods:
- Utilized expert-provided causal directed acyclic graphs (DAGs) to model dependencies within bags.
- Constructed states as dynamical Bayesian sufficient statistics of historical data to ensure Markovian transitions.
- Formulated the problem as a periodic MDP and generalized Bellman equations for online RL optimization.
- Evaluated the proposed method on mobile health trial data.
Main Results:
- The proposed state construction method ensures Markovian state transitions within and across bags.
- The developed online RL algorithm effectively handles non-stationarity within periodic MDPs.
- The constructed state was shown to achieve the maximal optimal value function for the periodic MDP.
- Successful evaluation on mobile health testbeds demonstrated practical applicability.
Conclusions:
- The novel RL framework successfully addresses non-Markovian and non-stationary dynamics in bagged decision time problems.
- The DAG-based state construction provides an effective way to manage complex temporal dependencies.
- The method offers a promising approach for optimizing sequential decision-making in domains like mobile health.
More Related Videos
07:05Operant Protocols for Assessing the Cost-benefit Analysis During Reinforced Decision Making by Rodents
Published on: September 10, 2018
07:42An Automated T-maze Based Apparatus and Protocol for Analyzing Delay- and Effort-based Decision Making in Free Moving Rodents
Published on: August 2, 2018
Related Concept Videos
Survival Tree
Building a Survival Tree
Constructing a...
Decision Making: Traditional Method
First, a specific claim about the population parameter is decided based on the research question and is stated in a simple form. Further, an opposing statement to this claim is also stated. These statements can act as null and alternative hypotheses, out of which a null hypothesis would be a...
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
Timing and Consequences on Behavior
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
Law of Effect
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...