Related Experiment Video
Updated: Aug 5, 2026

06:45
Automated Joint Space Detection Improves Bone Segmentation Accuracy
Published on: November 28, 2025
Explainable Two-Stage Xception-Swin Transformer Learning for Body-Part-Aware Fracture Detection in Musculoskeletal
Syed Baqir Hussain Shah1, Musfarah Wajid1, Syed Adil Hussain Shah2,3
1Department of Computer Science, COMSATS University Islamabad (CUI), Wah Campus, Wah 47000, Pakistan.
Journal of Imaging
|July 27, 2026
Summary
This study introduces a two-stage deep learning framework for improved upper-extremity X-ray analysis, enhancing fracture detection accuracy. The model effectively combines convolutional neural networks and transformers for better body-part recognition and abnormality identification.
Area of Science:
- Medical Imaging
- Artificial Intelligence
- Radiology
Background:
- Automated interpretation of upper-extremity radiographs is challenging due to subtle fractures and anatomical variations.
- Class imbalance in datasets further complicates accurate fracture detection in X-rays.
Purpose of the Study:
- To develop and evaluate a two-stage deep learning framework for improved body-part recognition and abnormality detection in musculoskeletal radiographs.
- To enhance the accuracy of automated fracture interpretation in upper-extremity X-rays.
Main Methods:
- A hybrid Xception-Swin deep learning model was developed, combining local structural features with transformer-based contextual features.
- The framework employed attention-based fusion and was evaluated in a two-stage process: body-part classification followed by abnormality detection within anatomical subsets.
- Performance metrics included accuracy, F1-score, AUC-ROC, Cohen's kappa, calibration, ablation studies, and zero-shot validation on the FracAtlas dataset.
Main Results:
- The model achieved high performance in body-part classification (accuracy=0.9643, macro F1=0.9574).
- Abnormality detection accuracy ranged from 0.7289 to 0.8538, with F1 scores from 0.7191 to 0.8508.
- Hybrid attention mechanisms improved performance, and preliminary cross-dataset generalizability was demonstrated via FracAtlas validation (AUC=0.8247, kappa=0.5812).
Conclusions:
- The proposed two-stage deep learning framework effectively improves automated interpretation of upper-extremity radiographs.
- Complementary fusion of Convolutional Neural Networks (CNNs) and Transformers shows promise for enhancing diagnostic accuracy in medical imaging.
- The study indicates preliminary cross-dataset generalizability, supporting the framework's potential for wider clinical application.
