Related Experiment Video
Updated: Jan 9, 2026

A Step-by-Step Implementation of DeepBehavior, Deep Learning Toolbox for Automated Behavior Analysis
Published on: February 6, 2020
Advances in deep reinforcement learning enable better predictions of human behavior in time-continuous tasks
Sabine Haberland1, Hannes Ruge1, Holger Frimmel1
1Institut of General Psychology, TUD Dresden University of Technology, Dresden, Germany.
Abstract:
Humans have to respond to everyday tasks with goal-directed actions in complex and time-continuous environments. However, modeling human behavior in such environments has been challenging. Deep Q-networks (DQNs), an application of deep learning used in reinforcement learning (RL), enable the investigation of how humans transform high-dimensional, time-continuous visual stimuli into appropriate motor responses. While recent advances in DQNs have led to significant performance improvements, it has remained unclear whether these advancements translate into improved modeling of human behavior. Here, we recorded motor responses in human participants (N = 23) while playing three distinct arcade games. We used stimulus features generated by a DQN as predictors for human data by fitting the DQN's response probabilities to human motor responses using a linear model. We hypothesized that advancements in RL models would lead to better prediction of human motor responses. Therefore, we used features from two recently developed DQN models (Ape-X and SEED) and a third baseline DQN to compare prediction accuracy. Compared to the baseline DQN, Ape-X and SEED involved additional structures, such as dueling and double Q-learning, and a long short-term memory, which considerably improved their performances when playing arcade games. Since the experimental tasks were time-continuous, we also analyzed the effect of temporal resolution on prediction accuracy by smoothing the model and human data to varying degrees. We found that all three models predict human behavior significantly above chance level. SEED, the most complex model, outperformed the others in prediction accuracy of human behavior across all three games. These results suggest that advances in deep RL can improve our capability to model human behavior in complex, time-continuous experimental tasks at a fine-grained temporal scale, thereby opening an interesting avenue for future research that complements the conventional experimental approach, characterized by its trial structure and use of low-dimensional stimuli.
Related Concept Videos
Reinforcement Schedules
Once a behavior is learned,...
Timing and Consequences on Behavior
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
Observational Learning
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Associative Learning
Classical conditioning, also known...
Law of Effect
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...

