Related Experiment Video
Updated: Jan 9, 2026

Qualitative and Quantitative Validation of Tools with Rating Scales Aimed at Assessing the Quality of University Service-Learning
Published on: August 29, 2025
From tetrachoric to kappa: How to assess reliability on binary scales.
1Department of Methodology and Statistics, Care and Public Health Research Institute (CAPHRI), Maastricht University, Maastricht, The Netherlands.
Estimating reliability for binary data is challenging. This study compares three methods, finding the normal approximation and frequentist approaches unreliable for binary scale reliability. The latent variable approach offers better insights.
Area of Science:
- Psychometrics
- Statistical Methods
Background:
- Reliability in psychometrics is vital for measurement accuracy.
- Binary outcomes pose unique challenges for reliability estimation compared to quantitative scales.
- Classical test theory and intraclass correlation coefficients are established for quantitative data.
Purpose of the Study:
- To review and link three major approaches for estimating reliability of single ratings on binary scales.
- To clarify conceptual relationships and evaluate performance across repeatability and reproducibility studies.
- To extend Bayesian methods for kappa coefficients and estimate manifest scale reliability.
Main Methods:
- Comparison of normal approximation, kappa coefficients, and latent variable approaches.
- Extension of Bayesian Dirichlet-multinomial method for multi-replicate kappa coefficients.
- Introduction of a Bayesian method for manifest scale reliability estimation from latent scale reliability.
Main Results:
- The normal approximation approach demonstrated poor performance.
- The frequentist approach exhibited unreliability due to singularity issues.
- The latent variable approach provides a robust framework for binary scale reliability.
Conclusions:
- The normal approximation and frequentist methods are not recommended for binary scale reliability.
- The latent variable approach offers a more reliable method for assessing binary outcome reliability.
- Refined practical recommendations are provided for reliability estimation in psychometric research.
More Related Videos
09:18Author Spotlight: Assessing the Reliability of Doppler Ultrasound in Measuring Leg Blood Flow
Published on: December 15, 2023
09:00Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
Published on: August 16, 2024
Related Concept Videos
Kendall's Coefficient of Concordance
Reliability and Validity
Kendall's Tau Test
A τ value of +1 indicates...
Spearman's Rank Correlation Test
Spearman's test calculates correlation by...
Introduction to Test of Independence
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
Ordinal Level of Measurement
Data measured using an ordinal scale are similar to nominal scale data, but there is one major difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks...