Related Experiment Video
Updated: May 22, 2025

Accuracy in Dental Medicine, A New Way to Measure Trueness and Precision
Published on: April 29, 2014
A model-free framework for evaluating the reliability of a new device with multiple imperfect reference standards
Ying Cui1, Qi Yu1, Amita Manatunga1
1Department of Biostatistics and Bioinformatics, Emory University, Atlanta, GA 30329, United States.
Evaluating new computer-aided diagnostic (CAD) devices is challenging without a gold standard. This study introduces a statistical framework to assess CAD device reliability using multiple imperfect reference standards, accounting for accuracy variations.
Area of Science:
- Medical Imaging
- Biostatistics
- Diagnostic Accuracy
Background:
- Evaluating computer-aided diagnostic (CAD) devices typically relies on gold standard tests.
- Gold standards are often unavailable in clinical studies, necessitating the use of multiple imperfect reference standards.
- Heterogeneity in diagnostic accuracy across reference standards can bias device evaluation.
Purpose of the Study:
- To develop a statistical framework for evaluating CAD devices using multiple imperfect reference standards.
- To address the challenge of heterogeneous diagnostic accuracy among reference standards.
- To provide an intuitive and easy-to-use method for device reliability assessment.
Main Methods:
- A statistical framework assessing agreement between a CAD device and a weighted sum of multiple imperfect reference standards.
- A model-free, unsupervised inductive procedure to determine weights for reference standards based on their relative reliability.
- Recursive weight assignment favoring more consistent and majority-opinion reference standards.
Main Results:
- The proposed framework effectively evaluates CAD devices by accounting for varying reference standard accuracy.
- Weights are assigned to reference standards without requiring modeling assumptions or external data.
- The method demonstrated its utility in evaluating a CAD device for kidney obstruction against multiple physician assessments.
Conclusions:
- The developed statistical framework offers a robust method for evaluating CAD devices when gold standards are absent.
- The unsupervised weighting procedure handles accuracy heterogeneity among reference standards effectively.
- This approach provides a reliable tool for assessing diagnostic device performance in real-world clinical settings.
Related Concept Videos
Uncertainty in Measurement: Accuracy and Precision
Systematic Error: Methodological and Sampling Errors
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Testing a Claim about Standard Deviation
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
Statistical Analysis: Overview
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
Random and Systematic Errors
Data Validation
Key parameters for method validation include:

