Related Experiment Video
Updated: May 24, 2025

04:23
A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
Published on: April 21, 2023
1.7K
Appropriateness of Thyroid Nodule Cancer Risk Assessment and Management Recommendations Provided by Large Language
1Radiological Sciences Department, College of Applied Medical Sciences, King Saud University, 11451, Riyadh, Saudi Arabia. mohalarifi@ksu.edu.sa.
Journal of Imaging Informatics in Medicine
|March 3, 2025
Summary
Large language models (LLMs) like ChatGPT, Gemini, and Claude show promise for thyroid nodule cancer risk assessment, aligning with clinical guidelines. While performance was similar, ChatGPT and Claude demonstrated slightly higher accuracy and clarity, indicating potential for clinical support with oversight.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Informatics
- Oncology Decision Support
Background:
- Large language models (LLMs) are increasingly explored for clinical applications.
- Evaluating the reliability of LLM recommendations in thyroid nodule cancer risk assessment is crucial for clinical adoption.
- Current clinical guidelines from the American Thyroid Association (ATA) and National Comprehensive Cancer Network (NCCN) provide a benchmark for assessment.
Purpose of the Study:
- To assess the appropriateness and reliability of thyroid nodule cancer risk assessment recommendations from ChatGPT, Gemini, and Claude.
- To compare the performance of these LLMs against established clinical guidelines (ATA and NCCN).
- To evaluate the readability and qualitative feedback on AI-generated responses.
Main Methods:
- Development of 24 clinically relevant questions based on ATA and NCCN guidelines.
- Evaluation of AI response readability using the Readability Scoring System.
- Assessment of 322 radiologists on AI-generated recommendations using quantitative (SPSS) and qualitative (Dedoose) analysis.
Main Results:
- No statistically significant differences in overall performance were found among ChatGPT, Gemini, and Claude.
- Claude (21.84) and ChatGPT (21.83) achieved slightly higher mean scores than Gemini (21.47).
- ChatGPT demonstrated the highest accuracy (92.5%) in appropriate responses, followed by Claude (92.1%) and Gemini (90.4%).
Conclusions:
- LLMs show potential for supporting thyroid nodule cancer risk assessment but require clinical oversight.
- Claude and ChatGPT exhibited comparable performance, with marginal differences in scores and accuracy.
- Further development is needed to enhance LLM reliability for widespread clinical use in oncology.

