Related Experiment Video
Updated: Aug 7, 2026

Development and Evaluation of 3D-Printed Cardiovascular Phantoms for Interventional Planning and Training
Published on: January 18, 2021
Evaluating artificial intelligence versus human-generated contours in the thorax
Oded Icht1,2,3, Patrik Farkas1, Ming Liu1,2
1Radiation Medicine Program, Princess Margaret Cancer Centre, University Health Network, Toronto, Ontario, Canada.
Background And Purpose:
Artificial intelligence (AI) auto-segmentation is increasingly used in radiotherapy to reduce contouring time, with physician review required before clinical use. These tools reliably achieve high geometric and dose agreement with manual organ of interest (OOI) contours, but whether such agreement predicts clinical acceptability on structured physician review is not well characterized.
Materials And Methods:
We analyzed twenty-three lung cancer cases using manual contours and AI-generated contours [Ethos-2 and RayStation (RS) 2023B] for the lungs, heart, esophagus, and spinal canal. We assessed quantitative agreement using Dice similarity coefficient (DSC), distance-to-agreement, and volumetric differences, as well as dose differences by comparing mean and maximum dose. For qualitative assessment, ten thoracic radiation oncologists each blindly reviewed ten cases, noting preferred and unacceptable contours, and differences were evaluated using Cochran's Q and pairwise McNemar tests.
Results:
Both AI platforms showed high geometric agreement with manual contours (DSC > 0.9 in over 70% of cases) and negligible dose differences across OOIs. However, reviewers more often found AI contours unacceptable for the esophagus (RS2023B) and the heart (Ethos-2). The most common reason for a contour being deemed unacceptable was insufficient anatomic accuracy rather than safety concerns.
Conclusions:
While AI-based auto-segmentation performed well on geometric and dose metrics, differences in physician acceptability persisted for some OOIs. These findings indicate that quantitative agreement is not a reliable predictor of clinical acceptability.

