Related Experiment Video
Updated: Jun 15, 2026

08:39
Longitudinal Micro-Computed Tomography Image Analysis for User-Defined Region of Interest in Critical-Sized Bone Defects
Published on: June 24, 2025
522
Benchmarking robustness of automated CT pancreas segmentation: achieving human-level reliability through
Felipe Oviedo1, Felipe Lopez-Ramirez2, Florent Tixier2
1AI for Good Lab, Microsoft Corporation, Redmond, WA, 98052, United States.
Radiology Advances
|December 15, 2025
Summary
Deep learning models for pancreas segmentation show promise but lack robustness. Active learning significantly improved reliability, achieving near-human performance and reducing workload.
Area of Science:
- Medical Imaging Analysis
- Artificial Intelligence in Healthcare
- Computational Anatomy
Background:
- Deep learning models for pancreas segmentation in CT scans have advanced rapidly.
- Current evaluation metrics (e.g., Dice, surface metrics) do not fully capture model robustness, defined as consistent human-level performance across diverse cases.
- Segmentation robustness is critical for clinical applications like early detection and quantitative biomarker analysis.
Purpose of the Study:
- To systematically evaluate the robustness of deep learning models for pancreas segmentation compared to human readers.
- To investigate the effectiveness of an active learning strategy for improving segmentation reliability.
Main Methods:
- Retrospective analysis of 903 CT scans, with 100 healthy test cases featuring 4 independent human segmentations each.
- Introduction of a Fractional Threshold (FT) metric to quantify robustness relative to human performance.
- Assessment of various deep learning models and implementation of an active learning approach for human-in-the-loop revision of uncertain predictions.
Main Results:
- The best 3D U-Net model achieved high overlap metrics (DSC: 0.88, NSD: 0.77), comparable to human readers (DSC: 0.89, NSD: 0.75).
- However, the Fractional Threshold (FT) metric revealed persistent model variability compared to human performance.
- Active learning with human-in-the-loop revision dramatically improved robustness (FT to 0.99) with minimal time investment per case (1.54 minutes), reducing workload by 23-fold.
Conclusions:
- Automated pancreas segmentation offers workload reduction but is limited by unpredictable failures in challenging cases.
- Active learning strategies are crucial for enhancing model reliability and bridging the performance gap between AI and human experts.
- Integrating active learning represents a significant step towards robust and clinically deployable AI for medical image segmentation.
Related Concept Videos
Imaging Studies for Cardiovascular System V: CT
Cardiac computed tomography (CT) scanning is an advanced cardiac imaging technique that utilizes CT technology, with or without intravenous (IV) contrast, to produce accurate cross-sectional virtual slices of specific areas of the heart, coronary circulation, and major blood vessels such as the aorta, pulmonary veins, and arteries. The computer processes these slices to generate three-dimensional images. Multidetector CT (MDCT) is a rapid form of CT scanning that captures multiple slices...
Multiple Comparison Tests
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...

