Conditional Reliability of Weighted Test Scores on a Bounded D-Scale
Dimiter M Dimitrov1,2, Dimitar V Atanasov3
1George Mason University, Fairfax, VA, USA.
Educational and Psychological Measurement
|December 23, 2025
Summary
This study extends conditional reliability to weighted scores on a bounded D-scale, offering new precision measures for item response theory. The research provides R syntax for calculating these important psychometric indicators.
Area of Science:
- Psychometrics
- Item Response Theory (IRT)
- Educational Measurement
Background:
- Previous research focused on conditional reliability for number-correct scores within Item Response Theory (IRT).
- These methods were conditioned on latent levels of the logit scale.
- Classical test theory (CTT) scores often lack reliability analysis across different latent trait levels.
Purpose of the Study:
- To investigate conditional reliability for classical-type weighted scores.
- To extend these concepts to a bounded scale, specifically the D-scale (0 to 1).
- To introduce new precision measures relevant to these scores.
Main Methods:
- Utilizing the D-scoring method for measurement on a bounded scale.
- Calculating conditional reliability for weighted D-scores.
- Developing and presenting conditional standard error, conditional signal-to-noise ratio, and marginal reliability.
Main Results:
- Conditional reliability was successfully extended to weighted scores on the D-scale.
- New precision measures were derived and presented alongside conditional reliability.
- The study provides practical R syntax for implementing these computations.
Conclusions:
- The D-scoring method framework allows for the analysis of conditional reliability in weighted scores.
- The introduced measures enhance the understanding of score precision across the D-scale.
- This work offers valuable tools for psychometricians and researchers in educational measurement.
Related Concept Videos
Testing a Claim about Standard Deviation
2.9K
A complete procedure to test a claim about population standard deviation or population variance is explained here.
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
2.9K
Reliability and Validity
13.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
13.7K
Wilcoxon Rank-Sum Test
661
The Wilcoxon rank-sum test, also known as the Mann-Whitney U test, is a nonparametric test used to determine if there is a significant difference between the distributions of two independent samples. This test is designed specifically for two independent populations and has the following key requirements:
661
z Scores and Area Under the Curve
18.2K
z scores are the standardized values obtained after converting a normal distribution into a standard normal distribution. A z score is measured in units of the standard deviation. The z score tells you how many standard deviations the value x is above (to the right of) or below (to the left of) the mean, μ. Values of x that are larger than the mean have positive z scores, and values of x that are smaller than the mean have negative z scores. If x equals the mean, then x has a z score of...
18.2K
Coefficient of Correlation
8.1K
The correlation coefficient, r, developed by Karl Pearson in the early 1900s, is numerical and provides a measure of strength and direction of the linear association between the independent variable x and the dependent variable y.
If you suspect a linear relationship between x and y, then r can measure how strong the linear relationship is.
What the VALUE of r tells us:
The value of r is always between –1 and +1: –1 ≤ r ≤ 1.
The size of the correlation r indicates the...
If you suspect a linear relationship between x and y, then r can measure how strong the linear relationship is.
What the VALUE of r tells us:
The value of r is always between –1 and +1: –1 ≤ r ≤ 1.
The size of the correlation r indicates the...
8.1K
Kendall's Coefficient of Concordance
911
Kendall's Coefficient of Concordance (W), also known as Kendall's W, is a non-parametric statistical measure used to assess the agreement or concordance between multiple raters or judges when they rank a set of items. It is often used when you have ordinal data (ranks) and you want to see if there is consistency or consensus among the raters. It is widely applied in research areas such as psychology, medicine, and social sciences, where multiple judges are asked to rank or rate subjects...
911


