在存在自我选择偏差和学习/实践效应的情况下估计测试-重试可靠性
William C M Belzak1, J R Lockwood1
1Duolingo, Pittsburgh, PA, USA.
Applied psychological measurement
|November 4, 2024
概括
从自选的重复器中估计测试重复测试可靠性可能会有偏见. 使用样本权重,回归和贝叶斯平均值的新方法提高了观察数据的准确性,这对于入学测试至关重要.
科学领域:
- 心理测量 心理测量 心理测量
- 统计建模 统计建模
背景情况:
- 测试重复测试的可靠性对于评估的有效性至关重要.
- 估计通常依赖于自我选择的测试重复者,引入潜在的偏见.
- 自我选择和时间变化的影响可能会扭曲可靠性指标.
研究的目的:
- 开发和评估从观测数据中进行无偏测试-重新测试可靠性估计的方法.
- 解决反复测试中自选和时间依赖效应引起的偏见.
- 提高可靠性估计的精度和通用性.
主要方法:
- 开发包括样本权重,多项式回归和贝叶斯模型平均值在内的方法.
- 这些方法应用于测试重复器的观察数据.
- 使用实证和模拟数据进行验证.
主要成果:
- 证明了测试-重新测试可靠性估计偏差的减少.
- 在可靠性估计的准确性方面取得了改进.
- 经验和模拟数据支持了拟议方法的有效性.
结论:
- 开发的方法有效地减轻了测试-重新测试可靠性估计中的偏差.
- 这些技术对于具有自我选择的重复器的设置是有价值的,例如入学测试.
- 这些方法可以概括为各种重复测量场景,并具有潜在的混因素.
相关概念视频
Hindsight Biases
3.4K
Hindsight bias leads you to believe that the event you just experienced was predictable, even though it really wasn’t. In other words, you knew all along that things would turn out the way they did. Can you relate this to the phrase "Hindsight is 20/20" now?
3.4K
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
Group Design
8.9K
The most basic experimental design involves two groups: the experimental group and the control group. The two groups are designed to be the same except for one difference— experimental manipulation. The experimental group gets the experimental manipulation—that is, the treatment or variable being tested—and the control group does not. Since experimental manipulation is the only difference between the experimental and control groups, we can be sure that any differences between...
8.9K
Surveys
14.7K
Often, psychologists develop surveys as a means of gathering data. Surveys are lists of questions to be answered by research participants, and can be delivered as paper-and-pencil questionnaires, administered electronically, or conducted verbally. Generally, the survey itself can be completed in a short time, and the ease of administering a survey makes it easy to collect data from a large number of people.
14.7K
Accuracy and Errors in Hypothesis Testing
176
Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
176
Blind Procedures
10.6K
Ideally, the people who observe and record the children’s behavior are unaware of who was assigned to the experimental or control group, in order to control for experimenter bias. Experimenter bias refers to the possibility that a researcher’s expectations might skew the results of the study. Remember, conducting an experiment requires a lot of planning, and the people involved in the research project have a vested interest in supporting their hypotheses. If the observers knew which...
10.6K


