Related Experiment Video
Updated: Jan 14, 2026

Accuracy in Dental Medicine, A New Way to Measure Trueness and Precision
Published on: April 29, 2014
Measuring Agreement in Diagnostics: A Practical Guide for Researchers
Sophie Vanbelle1, Christina Hernandez Engelhart2,3, Ellen Blix3
1Methodology and Statistics, CAPHRI, Maastricht University, Maastricht, Limburg, the Netherlands.
None:
Healthcare professionals routinely perform clinical examinations and diagnostic assessments. How the findings of these assessments are interpreted can have significant implications for patient care and outcomes. A recent systematic review on reliability and agreement studies in intrapartum fetal heart rate monitoring highlighted three methodological issues: (1) confusion between the concepts of agreement and reliability, (2) lack of clarity on how agreement and reliability measures are calculated when more than two raters are involved, and (3) confidence intervals seldom reported. This paper aims to clarify how agreement measures can be computed and interpreted when the outcome is binary (e.g., normal/abnormal test result). Using a motivating example in which five experienced obstetricians assessed 20 CTGs, we demonstrate how agreement can be defined, computed, and interpreted in various scenarios. The paper further explains the relationship between agreement measures and the concept of reliability, the distinction between intra- and inter-observer studies, and approaches to make statistical inference and sample size calculations. Particular emphasis is placed on the proportion of agreement, the proportion of specific agreement and kappa coefficients. A shiny application has also been developed to support researchers in their agreement studies. This work completes existing tools such as the Guidelines for Reporting Reliability and Agreement Studies (GRRAS), the Quality Appraisal Tool for Studies of Diagnostic Reliability (QAREL) and STARD guidelines for reporting diagnostic accuracy studies. It is intended to help researchers improve the methodological quality of studies that evaluate the agreement of clinical tests.
Related Concept Videos
Statistical Analysis: Overview
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
Receiver Operating Characteristic Plot
Accuracy and Precision
Variability: Analysis
The range is a simple measure of variability, indicating the difference between the highest and...
Uncertainty in Measurement: Reading Instruments

