Related Experiment Video
Updated: Jul 18, 2026

04:23
A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
Published on: April 21, 2023
Multimodal deep learning for predicting WHO/ISUP grading in renal tumors on CT using a self-attention-based model:
Takuma Usuzaki1,2, Eriya Matsuno1, Takashi Shizukuishi1,2
1Department of Diagnostic Radiology, Tohoku University Hospital, Sendai, Japan.
European Journal of Radiology Open
|July 12, 2026
Summary
The variable vision transformer (vViT) model accurately predicts renal tumor World Health Organization/International Society of Urological Pathology (WHO/ISUP) grade using multimodal data. Radiomic features were the most significant predictor for tumor grading.
Area of Science:
- Oncology
- Radiology
- Artificial Intelligence
Background:
- Renal tumors require accurate grading for prognosis and treatment.
- The World Health Organization/International Society of Urological Pathology (WHO/ISUP) system is standard for renal tumor classification.
- Integrating multimodal data offers potential for improved diagnostic accuracy.
Purpose of the Study:
- To evaluate the variable vision transformer (vViT) model's performance in predicting WHO/ISUP grade for renal tumors.
- To assess the contribution of clinical information, radiomic features, and CT images as model inputs.
- To identify the most influential factor in WHO/ISUP grade prediction.
Main Methods:
- Trained and validated the vViT model on a dataset of 111 patients (1398 images) and 15 patients (224 images).
- Classified renal tumors into low (WHO/ISUP grades 1-2) and high (grades 3-4) grades.
- Utilized permutation feature importance to determine the contribution of each input modality (clinical data, radiomics, CT images).
- Compared vViT performance against Vision Transformer (ViT), ConvNeXt, and ResNeXt models using the DeLong test.
Main Results:
- The vViT model achieved an accuracy of 0.811 and an AUC-ROC of 0.856 on the test dataset.
- vViT demonstrated superior performance compared to ViT and ResNeXt models (p < 0.05).
- Radiomic features were identified as the most dominant predictor for WHO/ISUP grade (p < 0.05).
Conclusions:
- The vViT model shows comparable performance to other models in predicting WHO/ISUP pathological grade using multimodal data.
- Radiomic analysis integrated with clinical data holds significant potential for enhancing renal tumor grading.
- Multimodal data integration, particularly radiomics, is promising for improving renal tumor classification accuracy.
