Related Experiment Video
Updated: Jul 17, 2025

06:48
Automated Segmentation of Cortical Grey Matter from T1-Weighted MRI Images
Published on: January 7, 2019
8.9K
USE-Evaluator: Performance metrics for medical image segmentation models supervised by uncertain, small or empty
Sophie Ostmeier1, Brian Axelrod1, Fabian Isensee2
1Stanford University, Center of Academic Medicine, 453 Quarry Rd, Palo Alto, CA 94304, United States of America.
Medical Image Analysis
|September 6, 2023
Summary
Standard performance metrics for medical image segmentation may not accurately reflect clinical utility, especially with uncertain or incomplete annotations. Rethinking evaluation is crucial for reliable AI in healthcare.
Area of Science:
- Medical image analysis
- Artificial intelligence in healthcare
- Machine learning model evaluation
Background:
- Traditional overlap metrics (e.g., Dice) are commonly used for medical image segmentation but may not align with clinical realities.
- Public datasets often have different case distributions and segmentation complexities than real-world clinical data.
- Existing metrics can be insensitive to challenges like uncertain, small, or empty reference annotations, impacting model development.
Purpose of the Study:
- To investigate the influence of uncertain, small, and empty reference annotations on performance metrics for medical image segmentation models.
- To identify suitable evaluation metrics that accurately reflect clinical value in challenging segmentation scenarios.
- To compare metric behavior across different datasets, including in-house stroke data and public datasets (BRATS 2019, Spinal Cord).
Main Methods:
- Analysis of metric performance on a stroke dataset with varying reference annotation qualities (uncertain, small, empty).
- Examination of a standard deep learning framework's predictions to understand metric behavior in challenging settings.
- Comparative analysis with results from the BRATS 2019 and Spinal Cord public datasets.
Main Results:
- Commonly used performance metrics can be misleading when reference annotations are uncertain, small, or empty.
- The study highlights a significant mismatch between standard evaluation practices and the demands of clinical application.
- Specific metrics require re-evaluation to ensure accurate assessment of segmentation model performance in diverse clinical contexts.
Conclusions:
- Rethinking the evaluation of medical image segmentation models is essential, particularly when dealing with imperfect or limited reference annotations.
- Current metrics may not adequately capture the clinical utility of segmentation models in real-world scenarios.
- The findings advocate for the development and adoption of more robust evaluation strategies for AI in medical imaging.

