Related Experiment Video
Updated: Jun 1, 2026

06:48
Automated Segmentation of Cortical Grey Matter from T1-Weighted MRI Images
Published on: January 7, 2019
9.0K
Quantitative Evaluation of Artificial Intelligence-Based Organ Segmentation Across Multiple Anatomic Sites Using 8
Lulin Yuan1, Quan Chen2, Hania Al-Hallaq3
1Department of Radiation Oncology, Virginia Commonwealth University, Richmond, Virginia.
Practical Radiation Oncology
|August 25, 2025
Summary
Commercial AI software shows significant variability in segmenting organs-at-risk (OARs), impacting clinical practice. Thorough testing and quality assurance are crucial for AI segmentation tools to ensure reliable patient care.
Area of Science:
- Medical Imaging Analysis
- Artificial Intelligence in Healthcare
- Radiotherapy Planning
Background:
- Automated segmentation of organs-at-risk (OARs) using artificial intelligence (AI) offers potential efficiency gains in radiotherapy planning.
- However, variability in AI software performance necessitates careful evaluation before clinical implementation.
Purpose of the Study:
- To assess the segmentation accuracy and variability of eight commercial AI software platforms across diverse anatomical sites.
- To compare AI-generated contours against clinical standards using multiple metrics.
- To provide recommendations for the clinical adoption of AI-based OAR segmentation.
Main Methods:
- Retrospective analysis of 160 planning CT datasets from head-and-neck, thorax, abdomen, and pelvis regions.
- Evaluation of 31 OARs segmented by AI software against clinical contours using Dice Similarity Coefficient (DSC), Hausdorff Distance (HD95), and relative added path length (RAPL).
- Statistical analysis using two-factor ANOVA to quantify inter-software and inter-patient variability.
Main Results:
- Significant inter-software and inter-patient variability in OAR segmentation accuracy was observed (p<0.05).
- Largest inter-software variations in DSC were noted for cervical esophagus (0.41), trachea (0.10), spinal cord (0.13), and prostate (0.17).
- Segmentation accuracy varied, with 7 OARs achieving mean DSC >0.9, 15 between 0.7-0.89, and others below 0.7. Over half (52%) of OARs showed RAPL < 0.1.
Conclusions:
- AI-based OAR segmentation software exhibits significant variability in performance across different platforms and patient anatomies.
- These findings underscore the critical need for rigorous software commissioning, validation, and ongoing quality assurance.
- Implementing standardized testing protocols is essential for safe and effective clinical integration of AI segmentation tools.

