Related Experiment Video
Updated: Aug 24, 2026

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
Published on: April 21, 2023
Comparative Performance of Multimodal Large Language Models in Grayscale Ultrasound-Based Classification of Thyroid
Ziman Chen1, Yingli Wang2, Fei Chen3
1Department of Health Technology and Informatics, The Hong Kong Polytechnic University, Kowloon, Hong Kong.
Background:
Multimodal large language models (LLMs) are increasingly being explored for medical image analysis, but their relative performance in thyroid ultrasound remains unclear.
Objective:
This study aimed to compare six publicly available multimodal LLMs for grayscale ultrasound-based classification of thyroid nodules.
Methods:
This prospective cross-sectional study included 178 patients with 239 thyroid nodules who underwent preoperative thyroid ultrasound followed by histopathological confirmation. Cropped grayscale ultrasound images of the maximal transverse and longitudinal views were analyzed by six publicly available multimodal LLMs: ChatGPT-5.4, Gemini 3.1 Pro, Claude Opus 4.6, Qwen3.6-Plus, Kimi K2.5, and ERNIE 5.0. All models were evaluated using the same image-input workflow and a standardized prompt, without fine-tuning or task-specific retraining. Agreement was assessed using Cohen's kappa, and diagnostic performance was evaluated using receiver operating characteristic (ROC) analysis. Radiologist benchmarks were included for comparison.
Results:
All six LLMs significantly distinguished benign from malignant nodules (all P ≤ 0.001). Gemini 3.1 Pro achieved the best overall performance, with a kappa value of 0.580 and an area under the ROC curve (AUC) of 77.1% (95% CI, 71.5%-82.7%). ChatGPT-5.4 and Qwen3.6-Plus each yielded an AUC of 73.5%, and Kimi K2.5 achieved an AUC of 71.3%. Claude Opus 4.6 and ERNIE 5.0 showed lower overall performance, with AUCs of 65.7% and 59.9%, respectively. The senior radiologist achieved higher diagnostic performance than all six LLMs.
Conclusion:
Publicly available multimodal LLMs showed measurable but heterogeneous performance in grayscale ultrasound-based thyroid nodule classification. Gemini 3.1 Pro demonstrated the best overall results, but none of the models matched senior radiologist-level performance.
