Related Experiment Video
Updated: May 22, 2026

In Vivo, Percutaneous, Needle Based, Optical Coherence Tomography of Renal Masses
Published on: March 30, 2015
Performance of a self-attention-based model in the task of differentiating clear cell renal cell carcinoma from other
Takuma Usuzaki1, Kengo Takahashi2, Hidenobu Takagi1,3
1Department of Diagnostic Radiology, Tohoku University Hospital, Sendai, Miyagi 980-8574, Japan.
Objectives:
To examine the performance of the variable Vision Transformer (vViT) in comparison with that of convolutional neural networks (CNNs) in the task of differentiating clear cell renal cell carcinoma (ccRCC) and non-ccRCC using CT images.
Methods:
The vViT was designed to use patient characteristics, radiomic features, and arterial phase CT images. The training and test datasets were constructed from the training set of the 2019 Kidney and Kidney Tumor Segmentation Challenge (C4KC-KiTS) dataset. The training dataset contained 153 patients with 1636 images (818 ccRCC, 818 non-ccRCC) and the test dataset contained 39 patients with 402 images (201 ccRCC, 201 non-ccRCC). After training, metrics including accuracy and the area under the curve of the receiver-operating characteristics (AUC-ROC) were calculated using the test dataset for vViT, Visual Geometry Group (VGG) 16, AlexNet, and GoogleNet in a patient-based approach. The metrics were calculated using CT images containing kidney and tumor. The AUC-ROC of vViT was compared with those of other models using the DeLong test.
Results:
vViT, VGG16, AlexNet, and GoogleNet achieved accuracies of 0.82 (95% CI, 0.72-0.86), 0.72 (0.62-0.78), 0.54 (0.45-0.62), and 0.61 (0.52-0.69), respectively. AUC-ROC values for vViT, VGG16, AlexNet, and GoogleNet were 0.91 (0.81-0.99), 0.70 (0.51-0.89), 0.53 (0.34-0.72), and 0.66 (0.46-0.85), respectively. The AUC-ROC of vViT was higher than that of VGG16, AlexNet, and GoogleNet (P < .05).
Conclusions:
Variable Vision Transformer achieved comparable performance to CNN using CT images in the task of differentiating ccRCC and non-ccRCC.
Advances In Knowledge:
We developed a new deep learning model termed vViT that simultaneously analyzes non-image information and medical images. Variable Vision Transformer can evaluate associations between input factors and outcome.