Related Experiment Video
Updated: Jan 29, 2026

Diaphragmatic Ultrasound in Adults: Image Acquisition and Interpretation
Published on: January 31, 2025
Can GPT-5.0 Interpret Thyroid Ultrasound Images? A Comparative TI-RADS Analysis with an Expert Radiologist
Yunus Yasar1, Sevde Nur Emir1, Muhammet Rasit Er1
1Department of Radiology, Umraniye Training and Research Hospital, University of Health Sciences, 34668 Istanbul, Turkey.
Large language models like GPT-5.0 show potential in interpreting thyroid ultrasounds using TI-RADS criteria, but struggle with specific features and tend to overestimate malignancy risk, limiting current clinical use.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Imaging Analysis
- Radiology
Background:
- Large language models (LLMs) are being explored for medical image interpretation, including thyroid ultrasounds (USs).
- Assessing thyroid nodules requires evaluating subtle echogenic and morphological details, a challenge for AI.
- The American College of Radiology (ACR) Thyroid Imaging Reporting and Data System (TI-RADS) provides a standardized framework for evaluating thyroid nodules.
Purpose of the Study:
- To evaluate GPT-5.0's performance in interpreting thyroid US images based on TI-RADS criteria.
- To compare GPT-5.0's assessments with those of an expert radiologist.
- To determine the diagnostic accuracy of GPT-5.0 for thyroid nodule malignancy risk stratification.
Main Methods:
- Retrospective analysis of 100 thyroid US cases with cytology-confirmed diagnoses (Bethesda II, V-VI).
- Independent assessment of TI-RADS features (composition, echogenicity, shape, margin, echogenic foci) by GPT-5.0 and a radiologist.
- Analysis of inter-rater agreement using Cohen's kappa and diagnostic performance metrics (sensitivity, specificity) with TR4-TR5 as positive for malignancy.
Main Results:
- GPT-5.0 showed substantial agreement with the radiologist for composition, shape, and margin, but only moderate for echogenicity and poor for echogenic foci.
- GPT-5.0 exhibited a tendency to up-classify nodules, assigning more TR4-TR5 categories, indicating a risk-averse bias.
- The radiologist achieved superior diagnostic performance (sensitivity 89%, specificity 85%) compared to GPT-5.0 (sensitivity 67%, specificity 49%), with GPT-5.0 frequently misclassifying benign nodules as suspicious.
Conclusions:
- GPT-5.0 can identify some TI-RADS features but struggles with microcalcifications and overestimates malignancy risk.
- The current performance limits GPT-5.0's standalone clinical utility for thyroid nodule risk stratification.
- Domain-specific training and adaptation may enhance LLMs for adjunctive roles in thyroid nodule assessment.
Related Concept Videos
Comparing the Survival Analysis of Two or More Groups
The Thyroid Gland
The follicles have a central cavity lined by simple cuboidal to squamous epithelial cells called follicular cells. These cells produce the glycoprotein...
Interpreting R Charts
An R chart plots the range of subsets of measurements collected from a process. Each point on the chart represents the range—defined as the difference between the maximum and minimum...
Interpreting Run Charts
Functions of Thyroid Hormones
TH is indispensable for the normal development and maturation of the skeletal, muscular, and nervous systems during fetal and childhood growth. It facilitates bone mineral turnover and regulates protein synthesis in developing tissues, contributing significantly to overall growth and...
Interpretation of Confidence Intervals
Confidence intervals have confidence coefficients that are crucial for their interpretation. The most common confidence coefficients are 0.90, 0.95, and 0.99, which can be written as percentages–90%, 95%, and 99%, respectively.
Suppose a person calculates a confidence interval with a confidence coefficient of 0.95. In that case, they can...

