Related Experiment Videos
TransUNet-GradCAM: a hybrid transformer-U-Net with self-attention and explainable visualizations for foot ulcer
Akwasi Asare1, Mary Sagoe2, Justice Williams Asare2
1Department of Computer Science, Faculty of Computing and Information Systems, Ghana Communication Technology University, Accra, Ghana. nasare34@yahoo.com.
Abstract:
Automated segmentation of diabetic foot ulcers (DFUs) supports clinical diagnosis, treatment planning, and longitudinal wound monitoring, but remains difficult owing to the heterogeneous appearance, irregular morphology, and cluttered backgrounds of ulcers in clinical photographs. Convolutional networks such as U-Net localise well but model long-range context poorly, whereas Vision Transformers capture global dependencies. We employ a hybrid ViT-bottleneck U-Net that combines a convolutional encoder-decoder with a Transformer bottleneck and attention-gated skip connections, and we emphasise that the contribution is a rigorously validated and explainable application rather than a new architecture. The model was trained on the public Foot Ulcer Segmentation Challenge (FUSeg) dataset with a hybrid Dice and cross-entropy loss, and all results are reported over five seeds as mean ± 95% confidence interval at a single fixed threshold. On the internal validation set it achieved a Dice of 0.8035 ± 0.0053 and an IoU of 0.7149 ± 0.0073 (HD95 = 19.74 px, ASSD = 6.12 px). A component-wise ablation showed that only the hybrid loss produced a statistically significant change in Dice (- 0.038, p < 0.001); the Transformer bottleneck, attention gates, and augmentation each had small, non-significant in-domain effects. External validation without retraining retained about 92% of internal Dice on the Advancing the Zenith of Healthcare (AZH) Wound Care Center cohort (Dice 0.7460, n = 278), while a small Medetec subset (n = 8) served only as a qualitative check, indicating partial rather than robust generalisation under domain shift. A quantitative explainability analysis found Grad-CAM more wound-localised (energy-in-mask 0.871 versus 0.102) but attention rollout significantly more faithful (p = 0.038, n = 200), showing the two are complementary. Predicted and expert wound areas agreed strongly (Pearson r = 0.944), and the model is lightweight (8.79 M parameters).