关于测试复试可靠性的不可靠性
1Institute of Psychology, Otto von Guericke University Magdeburg, Magdeburg, Germany.
Applied psychological measurement
|December 1, 2025
概括
测试复试系数 (TRC) 可能不可靠. 违反稳定的真分数和独立误差假设会影响TRC的估计,可能会在心理评估中误解测量可靠性.
科学领域:
- 心理测量 心理测量 心理测量
- 心理测量 心理测量
- 统计建模 统计建模
背景情况:
- 测试复试系数 (TRC) 是经典测试理论和心理评估可靠性的基石.
- TRC依赖于完全稳定的真得分和独立的错误得分的假设,这些得分通常未经测试.
- 违反这些核心假设可能会导致可靠性估计的重大偏差.
研究的目的:
- 调查违反TRC假设对可靠性估计的影响.
- 在不同条件下检查TRC性能,包括真得分稳定性和错误得分依赖性.
- 评估TRC在应用心理测量中的可解释性和适用性.
主要方法:
- 探索TRC假设的理论基础.
- 模拟研究使用人工数据来模拟不同的条件.
- 分析不同样本大小,真分数稳定性和错误分数依赖性的TRC性能.
主要成果:
- 降低真分数稳定性导致对可靠性的低估.
- 错误分数依赖可以人工膨胀TRC值,暗示虚假的可靠性.
- 违反假设的行为使得TRC的识别不足,从而损害了它的解释性.
结论:
- 当其基础假设不满足时,TRC的有效性是可疑的.
- 这些发现挑战了TRC在动态特征或不受控制的测量环境中的实用性.
- 当测量假设被违反时,其他可靠性指标可能更合适.
相关概念视频
Reliability and Validity
13.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
13.7K
Comparing Experimental Results: Student's t-Test
4.7K
The t-test is a statistical method used to compare the sample mean with a population mean or compare two means from two data sets. The test statistic is calculated from the standard deviation, mean, and number of measurements in the data set at a selected confidence interval and then compared to a table of critical values at this confidence level. If the test statistic is smaller than the critical value, the null hypothesis is accepted. In this case, we state that the difference between the...
4.7K
Uncertainty in Measurement: Accuracy and Precision
99.4K
Scientists typically make repeated measurements of a quantity to ensure the quality of their findings and to evaluate both the precision and the accuracy of their results. Measurements are said to be precise if they yield very similar results when repeated in the same manner. A measurement is considered accurate if it yields a result that is very close to the true or the accepted value. Precise values agree with each other; accurate values agree with a true value.
99.4K
Wilcoxon Rank-Sum Test
678
The Wilcoxon rank-sum test, also known as the Mann-Whitney U test, is a nonparametric test used to determine if there is a significant difference between the distributions of two independent samples. This test is designed specifically for two independent populations and has the following key requirements:
678
Accuracy and Errors in Hypothesis Testing
550
Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
550
Random and Systematic Errors
14.3K
Scientists always try their best to record measurements with the utmost accuracy and precision. However, sometimes errors do occur. These errors can be random or systematic. Random errors are observed due to the inconsistency or fluctuation in the measurement process, or variations in the quantity itself that is being measured. Such errors fluctuate from being greater than or less than the true value in repeated measurements. Consider a scientist measuring the length of an earthworm using a...
14.3K


