Related Experiment Video
Updated: Aug 14, 2026

Combining Reflectance Confocal Microscopy with Optical Coherence Tomography for Noninvasive Diagnosis of Skin Cancers via Image Acquisition
Published on: August 18, 2022
Interpretable multimodal fusion for skin lesion classification using dermoscopic images and patient metadata
Emrullah Sonuç1,2, Qusay Saihood2, Yusuf Yargı Baydilli3
1Computational Optimisation and Learning Lab, School of Computer Science, University of Nottingham, Nottingham, United Kingdom.
Introduction:
Accurate and early classification of skin lesions is essential for the early detection of disease and improved clinical outcomes. However, automated multiclass classification is still hindered by severe class imbalance, high inter-class visual similarity, and the comparatively under-explored use of complementary clinical metadata relative to image-only pipelines, despite growing interest in such metadata fusion. In line with the growing demand for artificial intelligence (AI) tools that integrate heterogeneous data in clinical scenarios, this study presents a multimodal machine learning-based computer-aided diagnosis (ML-CAD) framework that fuses dermoscopic images with patient metadata.
Methods:
The framework follows a five-phase pipeline. A structured multimodal balancing strategy combines class-wise Synthetic Minority Oversampling Technique for Nominal and Continuous features (SMOTENC) for metadata with controlled image augmentation. This is followed by cross-modal alignment to preserve clinical consistency. A fine-tuned Vision Transformer (ViT-B/16) at reduced input resolution performs feature extraction and late fusion with encoded metadata, yielding an 800-dimensional multimodal representation. Individual and stacked ensemble classifiers are subsequently trained on the fused features, with Bayesian optimization based on Gaussian process surrogates used for hyperparameter tuning. Gradient-weighted Class Activation Mapping (Grad-CAM) provides a visual interpretation across all seven lesion categories.
Results:
Experiments on the HAM10000 dataset demonstrate that the proposed SVM+KNN stacking ensemble achieved an accuracy of 98.55% and a ROC-AUC of 99.88%, achieving competitive accuracy among recently reported methods under broadly similar balanced settings.
Discussion:
The findings underscore the necessity of an AI framework that integrates imaging and clinical data to provide interpretable clinical decision support. This framework represents a step toward providing clinicians with multimodal, interpretable decision support. We emphasize, however, that external and prospective clinical validation remain necessary before translation into routine dermatological practice.