在大规模评估中检测项目不合适的强有力的方法.
Matthias von Davier1, Ummugul Bezirhan1
1Boston College, Chestnut Hill, MA, USA.
Educational and psychological measurement
|July 3, 2023
概括
本研究引入了一种新的方法来检测差异物品功能 (DIF),而不是假设完美的模型数据匹配. 它使用强大的异常值检测来识别不适当合适的项目,提高测量准确性.
科学领域:
- 心理测量 心理测量 心理测量
- 统计建模 统计建模
- 教育测量的教育测量.
背景情况:
- 精确的尺度构造需要识别项目不适合性和差异性项目功能 (DIF).
- 现有的方法通常假设完美的模型数据匹配,这可能是不现实的.
- 经典的测试理论和物品响应理论依赖于关于物品功能的明确假设.
研究的目的:
- 开发一种强大的方法来检测DIF,不需要完美的模型数据匹配.
- 为了提供一种更可靠的方法来评估物品适合规模结构.
主要方法:
- 利用了Tukey的受污染分布的概念.
- 采用了强大的异常值检测技术.
- 被标记的项目与模型数据不充分匹配.
主要成果:
- 成功识别了不适合模型数据的项目.
- 展示了对DIF检测的一种强有力的方法.
- 提供了一种不太依赖于理想化的统计假设的方法.
结论:
- 拟议的强大方法提高了DIF检测的准确性.
- 这种方法为尺度构造和测量提供了更实用的解决方案.
- 它解决了传统方法的局限性,因为它不假设完美的模型匹配.
相关概念视频
Quantifying and Rejecting Outliers: The Grubbs Test
1.7K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.7K
Detection of Gross Error: The Q Test
6.3K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.3K
Reliability and Validity
12.8K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.8K
Goodness-of-Fit Test
3.5K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
3.5K
Expected Frequencies in Goodness-of-Fit Tests
2.6K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.6K
Multiple Comparison Tests
3.9K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
3.9K


