Related Experiment Video
Updated: Oct 7, 2026

Computer-Aided Three-Dimensional Visualization in the Treatment of Locally Advanced Thyroid Cancer
Published on: June 9, 2023
Clinical evaluation of three deep learning contouring models for head and neck organs at risk and lymph node levels
Adam Miovecz1,2,3, Bettina Nagy3, Daniel Gugyeras3
1Department of Medical Imaging, Faculty of Health Sciences, University of Pecs, Kaposvar, 7400, Hungary.
Introduction:
Our aim was to compare the accuracy and clinical acceptability of head and neck organs at risk (OAR) and lymph node contours generated by three commercially available deep learning contouring (DLC) models: Mirada DLCExpert (Model D), Limbus AI (Model L), and MVision AI Contour+ (Model C).
Methods:
Seven OARs and cervical lymph node levels I to V were delineated by each model on twenty retrospectively selected patients with hypopharyngeal/laryngeal tumors (2020 to 2023). Automated contours were compared against expert reference contours using the Dice similarity coefficient (DSC), the 95th percentile Hausdorff distance (HD95), the median surface distance (MDSD) and the Jaccard index. Differences among the three models were tested with the Friedman test followed by pairwise Wilcoxon signed-rank tests with Holm-Bonferroni correction; agreement of structure volumes with the expert reference was tested with the Wilcoxon signed-rank test. A structured qualitative assessment classified each contour as requiring major corrections, minor corrections, or acceptable as is. A sensitivity analysis repeated every comparison using only reference contours independent of all three models.
Results:
Pooled over the seven OARs, mean DSC was 0.85 for Models L and C and 0.80 for Model D, but this ordering is produced entirely by the esophagus and the spinal canal, for which Model D applies a different anatomical definition (mean DSC 0.40 and 0.57, with model volumes 35% and 44% of the expert volume). Across the other five OARs Model D was the most accurate (0.92 against 0.84 and 0.85). After Holm-Bonferroni correction, Models L and C did not differ on any structure or metric. For lymph node levels no overlap metric separated the models in the primary analysis (all p > 0.05), though the omnibus test became significant in favour of Model D when restricted to reference contours independent of the models; Model D had the highest mean DSC (0.84) and the widest spread and Model C much the narrowest. Model C was the only model whose lymph node volumes showed no systematic difference from the expert reference (260.4 vs 256.6 cm3, mean difference + 3.8 cm3, 95% limits of agreement -46.5 to +54.1 cm3), against -33.4 and + 30.5 cm3 for Models D and L. Manual contouring averaged 28 min (OARs) and 16 min (lymph nodes); AI-assisted review took approximately 5 min per patient in total. In the subjective assessment Models L and C required fewer corrections than Model D (mean modification score 2.44 and 2.43 vs 2.03) and were comparable to each other.
Conclusion:
All three models produced contours of clinically usable quality apart from Model D on the esophagus and the spinal canal. Models L and C were not statistically different throughout. Model D differed from both on the esophagus and the spinal canal, where it applies a narrower anatomical definition, and was the most accurate model on the remaining five OARs, though that advantage was only partly confirmed when the analysis was restricted to reference contours independent of the models. Model C gave the most predictable lymph node output: the narrowest spread and the only volumes not differing significantly from the expert reference. Deep learning contouring models significantly improve head and neck treatment planning efficiency by providing clinically acceptable accuracy in a fraction of the time required for manual contouring.
