Related Experiment Videos
A severity-aware multi-task CNN-transformer framework for explainable diabetic foot ulcer triage and mobile
Abdullah1,2, Muhammad Ateeb Ather1,2, Zulaikha Fatima3
1Center for Computing Research, Instituto Politécnico Nacional, Mexico City, Mexico.
Introduction:
Diabetic foot ulcer (DFU) image classification can support timely clinical triage; however, many existing systems provide limited severity granularity, rely on single-source image datasets, and predominantly use pixel-level explanation methods. This study developed a seven-class, severity-aware, multi-task artificial intelligence framework for interpretable DFU image analysis and smartphone-class deployment.
Methods:
The framework combines domain-adaptive SimCLR pretraining, U-Net-based image refinement, EfficientNet-B0 feature extraction, a lightweight windowed transformer fusion block, and parallel classification and ordinal severity-regression heads. Four public data sources-DFUC 2020/2021, AZH, Medetec, and an IWGDF-aligned ordinal repository-were harmonized and deduplicated, yielding an analytical corpus of 24,925 original images. Evaluation included a 3,739-image held-out test set, five-fold stratified cross-validation, leave-one-source-out testing, an independent prospective clinician-audited cohort of 366 images, calibration and uncertainty analyses, concept-based interpretability assessment, and mobile-device profiling.
Results:
On the held-out test set, the framework achieved 95.5% accuracy (95% CI, 94.7-96.1) and a macro F1 score of 0.953, while mean five-fold cross-validation accuracy was 96.5% ± 0.3 percentage points. External ordinal severity estimation yielded a Spearman correlation of ρ = 0.91. In the clinician-audited cohort, weighted Cohen's κ for model-clinician severity agreement was 0.952. Model-assisted review was associated with a change in recorded management intent in 27.6% of cases, and 92.6% of generated explanations were rated as clinically useful. Full on-device inference, including gradient-weighted class activation mapping and concept-based explanation generation, required 322 ms per image on a Snapdragon 8 Gen 1 device.
Discussion:
The proposed framework demonstrates how multi-source representation learning, multi-task severity modeling, uncertainty assessment, concept-based interpretability, and efficient edge deployment can be integrated within a unified AI pipeline for DFU image analysis. The findings indicate high classification performance, strong ordinal severity agreement, clinically interpretable outputs, and technically feasible smartphone-class inference. Further multicenter prospective validation across diverse populations and acquisition conditions is required before translation into routine clinical decision-support settings.