Related Experiment Video
Updated: Mar 3, 2026

07:18
A Behavioral Test Battery for the Repeated Assessment of Motor Skills, Mood, and Cognition in Mice
Published on: March 2, 2019
20.1K
Test-Retest Reliability of Common Behavioral Decision Making Tasks.
Melissa T Buelow1, Wesley R Barnhart1
1Department of Psychology, The Ohio State University Newark, Newark, OH, USA.
Summary
The Balloon Analogue Risk Task (BART), Columbia Card Task (CCT), and Game of Dice Task (GDT) demonstrate moderate test-retest reliability for decision-making tasks. The Iowa Gambling Task (IGT) showed weak reliability, suggesting caution when using it for repeated assessments.
Area of Science:
- Cognitive Psychology
- Neuroscience
- Behavioral Economics
Background:
- Behavioral decision-making tasks are crucial for assessing risk preferences and decision-making under uncertainty.
- Understanding the test-retest reliability of these tasks is essential for their valid application in research and clinical settings.
- Previous research has yielded mixed findings regarding the reliability of commonly used decision-making tasks.
Purpose of the Study:
- To evaluate the test-retest reliability of four widely used behavioral decision-making tasks: the Iowa Gambling Task (IGT), Balloon Analogue Risk Task (BART), Columbia Card Task (CCT), and Game of Dice Task (GDT).
- To determine if performance on these tasks remains stable over a short interval.
- To inform the appropriate use of these tasks in longitudinal studies and clinical assessments.
Main Methods:
- Ninety-eight undergraduate students participated in the study.
- Participants completed two sessions of the IGT, BART, CCT, and GDT, separated by a three-week interval.
- Test-retest reliability was assessed using correlation coefficients and paired-samples t-tests.
Main Results:
- The BART, CCT, and GDT exhibited moderate test-retest reliability, indicating stable performance over the three-week period.
- The IGT demonstrated weak test-retest reliability, particularly in the initial trials (1-40), with only weak correlations for later trials (41-100).
- Significant differences in risk-taking behavior were observed across time for the IGT, GDT, and BART, suggesting potential practice or carryover effects.
Conclusions:
- The BART, CCT, and GDT are reliable for repeated measurements in decision-making research.
- The IGT shows limited reliability for assessing decision-making under risk, especially when administered repeatedly.
- Findings have implications for the design of studies employing repeated assessments of decision-making and risk behavior in both research and clinical contexts.

