On the Unreliability of Test-Retest Reliability
1Institute of Psychology, Otto von Guericke University Magdeburg, Magdeburg, Germany.
Applied Psychological Measurement
|December 1, 2025
Summary
The Test-Retest Coefficient (TRC) may be unreliable. Violating assumptions of stable true scores and independent errors biases TRC estimates, potentially misrepresenting measurement reliability in psychological assessments.
Area of Science:
- Psychometrics
- Psychological Measurement
- Statistical Modeling
Background:
- The Test-Retest Coefficient (TRC) is a cornerstone of reliability in Classical Test Theory and psychological assessments.
- TRC relies on assumptions of perfectly stable true scores and independent error scores, which are often untested.
- Violations of these core assumptions can lead to significant biases in reliability estimation.
Purpose of the Study:
- To investigate the impact of violated TRC assumptions on reliability estimates.
- To examine TRC performance under varying conditions of true score stability and error score dependence.
- To assess the interpretability and suitability of TRC in applied psychological measurement.
Main Methods:
- Exploration of the theoretical foundations of TRC assumptions.
- Simulation studies using artificial data to model varying conditions.
- Analysis of TRC performance across different sample sizes, true score stabilities, and error score dependencies.
Main Results:
- Decreased true score stability leads to underestimation of reliability.
- Error score dependence can artificially inflate TRC values, suggesting false reliability.
- Assumption violations render the TRC underidentified, compromising its interpretability.
Conclusions:
- The TRC's validity is questionable when its underlying assumptions are not met.
- Findings challenge the TRC's utility in dynamic traits or uncontrolled measurement contexts.
- Alternative reliability metrics may be more appropriate when measurement assumptions are violated.
Related Concept Videos
Reliability and Validity
13.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
13.7K
Comparing Experimental Results: Student's t-Test
4.7K
The t-test is a statistical method used to compare the sample mean with a population mean or compare two means from two data sets. The test statistic is calculated from the standard deviation, mean, and number of measurements in the data set at a selected confidence interval and then compared to a table of critical values at this confidence level. If the test statistic is smaller than the critical value, the null hypothesis is accepted. In this case, we state that the difference between the...
4.7K
Uncertainty in Measurement: Accuracy and Precision
99.4K
Scientists typically make repeated measurements of a quantity to ensure the quality of their findings and to evaluate both the precision and the accuracy of their results. Measurements are said to be precise if they yield very similar results when repeated in the same manner. A measurement is considered accurate if it yields a result that is very close to the true or the accepted value. Precise values agree with each other; accurate values agree with a true value.
99.4K
Wilcoxon Rank-Sum Test
678
The Wilcoxon rank-sum test, also known as the Mann-Whitney U test, is a nonparametric test used to determine if there is a significant difference between the distributions of two independent samples. This test is designed specifically for two independent populations and has the following key requirements:
678
Accuracy and Errors in Hypothesis Testing
550
Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
550
Random and Systematic Errors
14.3K
Scientists always try their best to record measurements with the utmost accuracy and precision. However, sometimes errors do occur. These errors can be random or systematic. Random errors are observed due to the inconsistency or fluctuation in the measurement process, or variations in the quantity itself that is being measured. Such errors fluctuate from being greater than or less than the true value in repeated measurements. Consider a scientist measuring the length of an earthworm using a...
14.3K


