Related Experiment Video
Updated: Jul 12, 2026

Longitudinal Micro-Computed Tomography Image Analysis for User-Defined Region of Interest in Critical-Sized Bone Defects
Published on: June 24, 2025
Variability in manual editing of head and neck organs of interest auto-segmentations: a multi-user, longitudinal
Rita Simões1, Sandra van der Velden1, Mark J Gooding2
1Department of Radiation Oncology, The Netherlands Cancer Institute, Amsterdam, the Netherlands.
Background:
Although auto-segmentation is widely used as a starting point for delineating organs of interest (OOIs), subsequent manual edits, and their variability over time and across users, may compromise workflow efficiency or treatment quality, limiting the intended benefits of auto-segmentation. Yet, longitudinal editing patterns are rarely assessed for clinical acceptability.
Aim:
To quantify editing of head-and-neck auto-segmentations over seven years, characterize variability among radiation therapy technologists (RTTs), and assess whether user-specific editing behavior aligns with perceived clinical acceptability.
Materials And Methods:
For 1051 head and neck patients treated since 2017, auto- and clinical segmentations of the parotid and submandibular glands, oral cavity and glottis were compared. Mean surface distance (MSD), normalized added path length (nAPL) and surface Dice Similarity coefficient (sDSC) were determined, and the locations of the most prominent edits were visualized, over time and by RTT. For a subset of recent cases, RTTs blindly rated the clinical acceptability of all participants' delineations.
Results:
MSD decreased for all structures except the glottis following the introduction of auto-segmentation. Switching from atlas-based to deep learning-based segmentation resulted in a marked nAPL reduction. Editing varied substantially over time and between RTTs. Delineations that were edited minimally or extensively were associated with the lowest clinical acceptability scores.
Conclusions:
Both excessive and minimal editing were associated with lower clinical acceptability, warranting the need for internal consensus to improve delineation consistency and avoid unnecessary workload or suboptimal contours. These findings underscore the importance of routine, per-user monitoring of auto-segmentation use in clinical practice.

