Related Experiment Video
Updated: Oct 26, 2025

Evaluation of a Point-of-Care Testing Analyzer for Measuring Peripheral Blood Leukocytes
Published on: March 22, 2022
Methods of assessing categorical agreement between correlated screening tests in clinical studies
Thomas J Zhou1, Sughra Raza2, Kerrie P Nelson1
1Department of Biostatistics, Boston University School of Public Health, Boston, MA, USA.
Abstract:
Advances in breast imaging and other screening tests have prompted studies to evaluate and compare the consistency between experts' ratings of existing with new screening tests. In clinical settings, medical experts make subjective assessments of screening test results such as mammograms. Consistency between experts' ratings is evaluated by measures of inter-rater agreement or association. However, conventional measures, such as Cohen's and Fleiss' kappas, are unable to be applied or may perform poorly when studies consist of many experts, unbalanced data, or dependencies between experts' ratings exist. Here we assess the performance of existing approaches including recently developed summary measures for assessing the agreement between experts' binary and ordinal ratings when patients undergo two screening procedures. Methods to assess consistency between repeated measurements by the same experts are also described. We present applications to three large-scale clinical screening studies. Properties of these agreement measures are illustrated via simulation studies. Generally, a model-based approach provides several advantages over alternative methods including the ability to flexibly incorporate various measurement scales (i.e. binary or ordinal), large numbers of experts and patients, sparse data, and robustness to prevalence of underlying disease.
More Related Videos
Related Concept Videos
Cochran's Q Test
Sign Test for Matched Pairs
To conduct the sign test, we first calculate the differences in...
Test for Homogeneity
Receiver Operating Characteristic Plot
Multiple Comparison Tests
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
McNemar's Test

