Related Experiment Videos
Multimodal deep learning integrating ultrasonographic and clinical data for enhanced rotator cuff tear identification
Ye Jiang1, Huining Xu1, Yuxuan Hu1
1Department of Ultrasound, Wuxi Ninth People's Hospital Affiliated to Soochow University, Wuxi, Jiangsu, China.
Background:
Ultrasonography (US) represents a practical and accessible imaging modality for rotator cuff tear (RCT) evaluation, yet diagnostic accuracy remains constrained by operator experience and the sonographic ambiguity of partial-thickness tears. Deep learning (DL) models applied to US imaging and machine learning models derived from clinical variables have each demonstrated independent promise, yet their integration within a unified diagnostic framework has not been systematically explored.
Objective:
This study aimed to develop and validate a multimodal DL model integrating US imaging with structured clinical examination data for RCT detection.
Methods:
A retrospective cohort of 947 patients presenting with shoulder pain who underwent both shoulder US and MRI was partitioned into training and testing sets at a 7:3 ratio. Patients with MRI-confirmed supraspinatus RCT were designated RCT-positive, while those with supraspinatus tendinopathy or subacromial bursitis were designated RCT-negative. Four model configurations were developed: a clinical-only model, a single-view US model, a dual-view US model, and a full multimodal model (MM) combining dual-view US with clinical variables through feature-level fusion. All models were trained using five-fold cross-validation and evaluated on the testing set for discriminative performance, calibration, and clinical utility. Model interpretability was assessed using Gradient-weighted Class Activation Mapping (Grad-CAM) for the image encoder and SHapley Additive exPlanations (SHAP) for the clinical encoder.
Results:
Among 947 patients, 370 (39.1%) were RCT-positive and 577 (60.9%) were RCT-negative. MM achieved the highest area under the receiver operating characteristic curve (AUROC) of 0.935, with sensitivity of 0.946, specificity of 0.786, and F1 score of 0.830, significantly outperforming all unimodal configurations. Decision curve analysis confirmed superior net benefit of MM across clinically relevant threshold probabilities. Grad-CAM heatmaps demonstrated anatomically coherent attention within the supraspinatus tendon, while SHAP analysis identified the drop arm test, external rotation lag sign, and empty can test as the dominant clinical predictors.
Conclusion:
The proposed multimodal DL framework offers a validated and interpretable tool for MRI-independent RCT diagnosis, with potential applicability in resource-constrained clinical settings.