Related Experiment Videos
An Attention-Enhanced Multimodal Hybrid Model for Skin Cancer Diagnosis Using Imaging and Clinical Data
Fatima Erik Dogan1, Merve Kesim Onal2, Harun Bingol3
1Department of Dermatology, Elazig Fethi Sekin City Hospital, Elazig 23300, Türkiye.
Abstract:
Background/Objectives: Skin cancer is one of the most common diseases worldwide, with a high mortality rate. Due to its ability to metastasize, the disease can progress to more serious stages over time. This article proposes a hybrid model based on feature engineering that will play a critical role in the early diagnosis of the disease. Methods: The developed model in this paper utilizes the well-known Vision Transformer (ViT) and Convolutional Neural Network (CNN) models for feature extraction from images in the dataset, while the FT-Transformer, Excel Former, SAINT, GRANDE, PTaRL, and TabTransformer architectures are used for feature extraction from clinical data. Furthermore, this study was developed using a very large pool of classifiers, including 13 classifiers. Fine-tuning was applied to improve the performance of the developed model. Channel attention mechanisms were incorporated into the study to ensure that the proposed model focuses on the diseased area. The PAD-UFES-20 dataset was used during the experiments. Class weighting was applied to the proposed model to prevent class-based imbalance in the PAD-UFES-20 dataset. Results: Six distinct CNN and four distinct ViT models were compared to the developed model. The developed model achieved a highly competitive Area Under the Curve (AUC) rate of 96.41%. The study was conducted using a dataset containing both clinical and imaging data. Conclusions: The proposed model is thought to help dermatologists diagnose skin cancer.