Related Experiment Video
Updated: Jan 10, 2026

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
Comparative Evaluation of Deep Learning Models for 3D Segmentation and Volumetry of Vestibular Schwannomas Using
Parv M Mehta1,2, Sahika B Yayli1, Pranjal Rai3
1From the Department of Radiology (P.M.M., S.B.Y., D.J.B., V.M.S., J.C.B., B.J.E., G.B.), Mayo Clinic, Rochester, Minnesota.
Background And Purpose:
3D segmentation and volumetry of vestibular schwannomas (VSs) is a more accurate method to determine tumor growth on serial imaging, but manual annotation is time-consuming to implement in routine clinical practice. We evaluated and compared 5 deep learning-based segmentation models (nnUNet [Base, ResEncL], U-Mamba, UNETR, and MedSAM) for 3D VS segmentation and volumetry, and we examined the robustness to acquisition heterogeneity and generalization on an external cohort.
Materials And Methods:
Our refined internal data set consisted of T1-contrast-enhanced images, including 2692 scans (n =383 patients) for training and 277 scans (n = 97 patients) for testing. Post-model training and validation, performance was evaluated on both internal and a publicly available external test set (n = 241) using the Dice similarity coefficient, maximum distance between the predicted and ground truth boundaries (Hausdorff distance), surface-to-surface (S2S) distance, and relative volume error (RVE). A subanalysis of the model performance was also performed to evaluate the impact of tumor volumes and data set heterogeneity.
Results:
The median Dice score on the external test set varied between 0.899 and 0.927 with U-Mamba achieving the highest performance, followed by nnUNet (Base and ResEncL). For these top 3 models, the median Hausdorff distance was 3.59 mm, while the 95th percentile Hausdorff distance was 1.6 mm. The S2S distance was <1 mm, and the median RVE (%) varied between 0.07 and 0.08. The median Dice scores were lower, 0.848-0.85, for smaller tumors (<200 mm3) and higher for tumors of >400 mm3 (median Dice score, 0.925-0.932).
Conclusions:
Models based on convolutional neural networks, transformer networks, and foundational models show robust performance for VS segmentation. Given the consistently high performance and self-optimizing frameworks of convolutional neural network-based models (U-Mamba, nnUNet), these may be more suitable for clinical applications.
