在不需要参考标准的情况下,统计测试人工智能细分设备和多专家的人类小组之间的基于重叠的性能协议,而不需要参考标准
Tingting Hu1, Berkman Sahiner1, Shuyue Guan1
1U.S. Food and Drug Administration, Silver Spring, Maryland, United States.
Journal of medical imaging (Bellingham, Wash.)
|October 24, 2025
概括
一种新的统计方法通过比较人工智能对专家和专家对专家的性能来评估人工智能 (AI) 医疗成像细分. 这种配对测试方法评估了人工智能与人类专家的协议,而不需要参考标准.
科学领域:
- 医疗成像医学成像
- 人工智能的人工智能
- 统计分析 统计分析
背景情况:
- 基于人工智能的医学成像设备通常执行病变或器官细分.
- 目前的评估方法依赖于聚合的参考标准和像Dice系数这样的指标,这些指标也有局限性.
- 定义有意义的成功标准和建立AI细分评估的黄金标准仍然具有挑战性.
研究的目的:
- 开发一种新的统计方法来评估AI细分性能.
- 评估人工智能设备与多位人类专家之间的协议.
- 克服现有评估方法的局限性,需要参考标准.
主要方法:
- 提出了一种对联测试的统计方法,将AI细分性能与多个人类专家进行比较.
- 该方法评估了AI与专家之间的差异相对于专家与专家之间的差异.
- 使用统计和基于图像的模拟进行验证,并应用于来自肺图像数据库联盟的肺图像的AI细分.
主要成果:
- 统计模拟证明了对I型和II型错误的有效控制.
- 基于图像的模拟显示了在评估协议方面可接受的性能.
- 该方法成功地应用于现实世界的医学成像数据.
结论:
- 引入了用于AI细分评估的新配对测试统计方法.
- 该方法可以在没有参考标准的情况下评估人工智能与人类专家的协议.
- 为验证AI医疗成像设备提供了有价值的工具.
相关概念视频
Multiple Comparison Tests
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
Introduction to Nonparametric Statistics
Nonparametric statistics offer a powerful alternative to traditional parametric methods, useful when assumptions about the population distribution cannot be made. Unlike parametric tests, which require data to follow a specific distribution with well-defined parameters (such as the mean and standard deviation), nonparametric tests do not require such constraints. This makes them particularly valuable when dealing with small sample sizes, skewed data, or ordinal and categorical variables.
One of...
One of...
Measures of Intelligence
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this; it...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this; it...


