Related Experiment Video
Updated: Aug 14, 2026

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
Quantifying User Satisfaction: Weighted Metric Approach for Evaluating Deep Learning-Based Thigh MRI Segmentations
Falko Ensle1, Ilker Özgür Koska2,3, Nina Derron4
1Diagnostic and Interventional Radiology, University Hospital of Zurich, 8091 Zurich, Switzerland.
Abstract:
Background: The aim of this paper is to develop a clinically useful benchmark signature of deep learning (DL)-based MRI segmentations and objectively predict user satisfaction based on a weighted combination of multiple performance metrics. Methods: This post hoc analysis of a prospective study analyzed MRI data from 68 patients (29.4 ± 5.6 years; 45% female) acquired during a randomized clinical trial. Fat fraction maps of axial Dixon MRI were used to segment six classes of the thigh. Two radiologists qualitatively scored DL-based segmentations using a 5-point Likert scale. A hybrid deep learning model (MobileNetV2 and DINOv2) was developed to predict Likert scores directly from MRI slices and segmentation maps. Additionally, a weighted combination of quantitative metrics (Dice score, Hausdorff distance, Jaccard index) was determined to best correlate with the Likert scores. Statistical analyses included Cohen's kappa, Spearman correlation, and regression model performance (MAE, RMSE, ROC_AUC). Results: Interreader agreement for Likert scores was near-perfect (κ = 0.82). The DINOv2 model achieved lower MAE (0.2343) and RMSE (0.1129) than MobileNetV2 (0.3488 and 0.431, respectively) in predicting Likert scores. The weighted metric combination model, using Lasso regression, predicted Likert = 5 with 84.1% accuracy, 100% recall, and ROC_AUC of 0.824. Individual metrics showed weak-to-moderate correlation with Likert scores, with Dice and Jaccard indices for extensor intramuscular fat exhibiting the highest correlation (r = 0.37 and r = 0.34, respectively). Conclusions: The proposed DL-based model accurately predicted radiologist-assigned Likert five scores, while the weighted metric combination model aligned closely with user satisfaction, offering a practical alternative to manual validation for clinical adoption of automated segmentation tools.