Statistical testing of agreement in overlap-based performance between an AI segmentation device and a multi-expert
Tingting Hu1, Berkman Sahiner1, Shuyue Guan1
1U.S. Food and Drug Administration, Silver Spring, Maryland, United States.
Journal of Medical Imaging (Bellingham, Wash.)
|October 24, 2025
Summary
A new statistical method evaluates artificial intelligence (AI) medical imaging segmentation by comparing AI-to-expert and expert-to-expert performance. This paired-testing approach assesses AI agreement with human experts without needing a reference standard.
Area of Science:
- Medical Imaging
- Artificial Intelligence
- Statistical Analysis
Background:
- AI-based medical imaging devices commonly perform lesion or organ segmentation.
- Current evaluation methods rely on aggregated reference standards and metrics like Dice coefficient, which have limitations.
- Defining meaningful success criteria and establishing a gold standard for AI segmentation evaluation remain challenging.
Purpose of the Study:
- To develop a novel statistical method for evaluating AI segmentation performance.
- To assess the agreement between AI devices and multiple human experts.
- To overcome the limitations of existing evaluation methods that require a reference standard.
Main Methods:
- Proposed a paired-testing statistical method to compare AI segmentation performance against multiple human experts.
- The method evaluates AI-to-expert dissimilarity relative to expert-to-expert dissimilarity.
- Validated using statistical and image-based simulations, and applied to AI segmentation of lung images from the Lung Image Database Consortium.
Main Results:
- Statistical simulations demonstrated effective control of Type I and Type II errors.
- Image-based simulations showed acceptable performance in assessing agreement.
- The method was successfully applied to real-world medical imaging data.
Conclusions:
- Introduced a new paired-testing statistical method for AI segmentation evaluation.
- The method enables assessment of AI agreement with human experts without a reference standard.
- Provides a valuable tool for validating AI medical imaging devices.
Related Concept Videos
Multiple Comparison Tests
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
Introduction to Nonparametric Statistics
Nonparametric statistics offer a powerful alternative to traditional parametric methods, useful when assumptions about the population distribution cannot be made. Unlike parametric tests, which require data to follow a specific distribution with well-defined parameters (such as the mean and standard deviation), nonparametric tests do not require such constraints. This makes them particularly valuable when dealing with small sample sizes, skewed data, or ordinal and categorical variables.
One of...
One of...
Measures of Intelligence
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this; it...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this; it...


