境界付きDスケールにおける加重テストスコアの条件付き信頼性
Dimiter M Dimitrov1,2, Dimitar V Atanasov3
1George Mason University, Fairfax, VA, USA.
Educational and psychological measurement
|December 23, 2025
まとめ
本研究は、条件付き信頼性を境界付きDスケール上の加重スコアに拡張し、項目応答理論のための新しい精度指標を提供する。この研究は、これらの重要な心理測定指標を計算するためのR構文を提供する。
科学分野:
- 心理測定学
- 項目応答理論(IRT)
- 教育測定
背景:
- 以前の研究は、項目応答理論(IRT)における正答数スコアの条件付き信頼性に焦点を当てていました。
- これらの方法は、ロジットスケールの潜在レベルに条件付けられていました。
- 古典的テスト理論(CTT)のスコアは、異なる潜在特性レベルにわたる信頼性分析を欠いていることがよくあります。
研究 の 目的:
- 古典的な加重スコアの条件付き信頼性を調査すること。
- これらの概念を境界付きスケール、特にDスケール(0から1)に拡張すること。
- これらのスコアに関連する新しい精度指標を導入すること。
主な方法:
- 境界付きスケールでの測定のためのDスコアリング法を利用すること。
- 加重Dスコアの条件付き信頼性を計算すること。
- 条件付き標準誤差、条件付き信号対雑音比、および周辺信頼性を開発および提示すること。
主要な成果:
- 条件付き信頼性は、Dスケール上の加重スコアに正常に拡張されました。
- 新しい精度指標が導出され、条件付き信頼性と並んで提示されました。
- この研究は、これらの計算を実装するための実践的なR構文を提供します。
結論:
- Dスコアリング法の枠組みは、加重スコアにおける条件付き信頼性の分析を可能にします。
- 導入された指標は、Dスケール全体にわたるスコア精度 の理解を深めます。
- この研究は、心理測定学者および教育測定の研究者にとって価値のあるツールを提供します。
関連する概念動画
Testing a Claim about Standard Deviation
2.9K
A complete procedure to test a claim about population standard deviation or population variance is explained here.
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
2.9K
Reliability and Validity
13.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
13.7K
Wilcoxon Rank-Sum Test
661
The Wilcoxon rank-sum test, also known as the Mann-Whitney U test, is a nonparametric test used to determine if there is a significant difference between the distributions of two independent samples. This test is designed specifically for two independent populations and has the following key requirements:
661
z Scores and Area Under the Curve
18.2K
z scores are the standardized values obtained after converting a normal distribution into a standard normal distribution. A z score is measured in units of the standard deviation. The z score tells you how many standard deviations the value x is above (to the right of) or below (to the left of) the mean, μ. Values of x that are larger than the mean have positive z scores, and values of x that are smaller than the mean have negative z scores. If x equals the mean, then x has a z score of...
18.2K
Coefficient of Correlation
8.1K
The correlation coefficient, r, developed by Karl Pearson in the early 1900s, is numerical and provides a measure of strength and direction of the linear association between the independent variable x and the dependent variable y.
If you suspect a linear relationship between x and y, then r can measure how strong the linear relationship is.
What the VALUE of r tells us:
The value of r is always between –1 and +1: –1 ≤ r ≤ 1.
The size of the correlation r indicates the...
If you suspect a linear relationship between x and y, then r can measure how strong the linear relationship is.
What the VALUE of r tells us:
The value of r is always between –1 and +1: –1 ≤ r ≤ 1.
The size of the correlation r indicates the...
8.1K
Kendall's Coefficient of Concordance
911
Kendall's Coefficient of Concordance (W), also known as Kendall's W, is a non-parametric statistical measure used to assess the agreement or concordance between multiple raters or judges when they rank a set of items. It is often used when you have ordinal data (ranks) and you want to see if there is consistency or consensus among the raters. It is widely applied in research areas such as psychology, medicine, and social sciences, where multiple judges are asked to rank or rate subjects...
911


