Estimating Test-Retest Reliability in the Presence of Self-Selection Bias and Learning/Practice Effects.
William C M Belzak1, J R Lockwood1
1Duolingo, Pittsburgh, PA, USA.
Applied Psychological Measurement
|November 4, 2024
Summary
Estimating test-retest reliability from self-selected repeaters can be biased. New methods using sample weighting, regression, and Bayesian averaging improve accuracy for observational data, crucial for admissions testing.
Area of Science:
- Psychometrics
- Statistical modeling
Background:
- Test-retest reliability is crucial for assessment validity.
- Estimates often rely on self-selected test repeaters, introducing potential bias.
- Self-selection and time-varying effects can distort reliability measures.
Purpose of the Study:
- To develop and evaluate methods for unbiased test-retest reliability estimation from observational data.
- To address biases arising from self-selection and time-dependent effects in repeated testing.
- To improve the precision and generalizability of reliability estimates.
Main Methods:
- Development of methods including sample weighting, polynomial regression, and Bayesian model averaging.
- Application of these methods to observational data from test repeaters.
- Validation using both empirical and simulated data.
Main Results:
- Demonstrated reduction in bias for test-retest reliability estimates.
- Showcased improvements in the precision of reliability estimations.
- Empirical and simulated data supported the effectiveness of the proposed methods.
Conclusions:
- The developed methods effectively mitigate bias in test-retest reliability estimation.
- These techniques are valuable for settings with self-selected repeaters, such as admissions testing.
- The methods generalize to various repeated measurement scenarios with potential confounding factors.
Keywords:
Bayesian model averagingentropy balancingminimum discriminant information adjustmentpolynomial regressionself-selection biastest-retest reliabilityMore Related Videos
Related Concept Videos
Hindsight Biases
3.4K
Hindsight bias leads you to believe that the event you just experienced was predictable, even though it really wasn’t. In other words, you knew all along that things would turn out the way they did. Can you relate this to the phrase "Hindsight is 20/20" now?
3.4K
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
Group Design
8.9K
The most basic experimental design involves two groups: the experimental group and the control group. The two groups are designed to be the same except for one difference— experimental manipulation. The experimental group gets the experimental manipulation—that is, the treatment or variable being tested—and the control group does not. Since experimental manipulation is the only difference between the experimental and control groups, we can be sure that any differences between...
8.9K
Surveys
14.7K
Often, psychologists develop surveys as a means of gathering data. Surveys are lists of questions to be answered by research participants, and can be delivered as paper-and-pencil questionnaires, administered electronically, or conducted verbally. Generally, the survey itself can be completed in a short time, and the ease of administering a survey makes it easy to collect data from a large number of people.
14.7K
Accuracy and Errors in Hypothesis Testing
176
Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
176
Blind Procedures
10.6K
Ideally, the people who observe and record the children’s behavior are unaware of who was assigned to the experimental or control group, in order to control for experimenter bias. Experimenter bias refers to the possibility that a researcher’s expectations might skew the results of the study. Remember, conducting an experiment requires a lot of planning, and the people involved in the research project have a vested interest in supporting their hypotheses. If the observers knew which...
10.6K


