权重测试成绩在有限的D尺度上的条件可靠性
Dimiter M Dimitrov1,2, Dimitar V Atanasov3
1George Mason University, Fairfax, VA, USA.
Educational and psychological measurement
|December 23, 2025
概括
这项研究将条件可靠性扩展到边界D尺度上的加权得分,为项目响应理论提供了新的精度指标. 这项研究提供了R语法来计算这些重要的心理指标.
科学领域:
- 心理测量 心理测量 心理测量
- 项目响应理论 (IRT)
- 教育测量教育的测量
背景情况:
- 之前的研究集中在物件响应理论 (IRT) 中对正确数值得分的条件可靠性.
- 这些方法取决于逻辑尺度的潜在水平.
- 经典测试理论 (CTT) 评分往往缺乏跨不同潜伏特征水平的可靠性分析.
研究的目的:
- 调查经典类型加权分数的条件可靠性.
- 将这些概念扩展到一个有界的尺度,特别是D尺度 (0到1).
- 引入与这些分数相关的新精度指标.
主要方法:
- 使用D-评分方法在边界尺度上进行测量.
- 计算加权D分数的条件可靠性.
- 开发和呈现有条件的标准误差,有条件的信号噪声比和边际可靠性.
主要成果:
- 条件可靠性成功地扩展到D级别的加权分数.
- 导出了新的精度指标,并与条件可靠性一起呈现.
- 该研究为实现这些计算提供了实用的R语法.
结论:
- D-评分方法框架允许在加权分数中分析条件可靠性.
- 引入的措施有助于人们更好地了解D级的得分精度.
- 这项工作为心理测量学家和教育测量研究人员提供了有价值的工具.
相关概念视频
Testing a Claim about Standard Deviation
2.9K
A complete procedure to test a claim about population standard deviation or population variance is explained here.
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
2.9K
Reliability and Validity
13.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
13.7K
Wilcoxon Rank-Sum Test
661
The Wilcoxon rank-sum test, also known as the Mann-Whitney U test, is a nonparametric test used to determine if there is a significant difference between the distributions of two independent samples. This test is designed specifically for two independent populations and has the following key requirements:
661
z Scores and Area Under the Curve
18.2K
z scores are the standardized values obtained after converting a normal distribution into a standard normal distribution. A z score is measured in units of the standard deviation. The z score tells you how many standard deviations the value x is above (to the right of) or below (to the left of) the mean, μ. Values of x that are larger than the mean have positive z scores, and values of x that are smaller than the mean have negative z scores. If x equals the mean, then x has a z score of...
18.2K
Coefficient of Correlation
8.1K
The correlation coefficient, r, developed by Karl Pearson in the early 1900s, is numerical and provides a measure of strength and direction of the linear association between the independent variable x and the dependent variable y.
If you suspect a linear relationship between x and y, then r can measure how strong the linear relationship is.
What the VALUE of r tells us:
The value of r is always between –1 and +1: –1 ≤ r ≤ 1.
The size of the correlation r indicates the...
If you suspect a linear relationship between x and y, then r can measure how strong the linear relationship is.
What the VALUE of r tells us:
The value of r is always between –1 and +1: –1 ≤ r ≤ 1.
The size of the correlation r indicates the...
8.1K
Kendall's Coefficient of Concordance
911
Kendall's Coefficient of Concordance (W), also known as Kendall's W, is a non-parametric statistical measure used to assess the agreement or concordance between multiple raters or judges when they rank a set of items. It is often used when you have ordinal data (ranks) and you want to see if there is consistency or consensus among the raters. It is widely applied in research areas such as psychology, medicine, and social sciences, where multiple judges are asked to rank or rate subjects...
911


