Related Experiment Video
Updated: Jan 29, 2026

Diaphragmatic Ultrasound in Adults: Image Acquisition and Interpretation
Published on: January 31, 2025
Can GPT-5.0 Interpret Thyroid Ultrasound Images? A Comparative TI-RADS Analysis with an Expert Radiologist
Yunus Yasar1, Sevde Nur Emir1, Muhammet Rasit Er1
1Department of Radiology, Umraniye Training and Research Hospital, University of Health Sciences, 34668 Istanbul, Turkey.
Abstract:
Background/Objectives: Multimodal large language models (LLMs) may directly interpret medical images, including thyroid ultrasounds (USs). Whether these models can reliably assess thyroid nodules-where subtle echogenic and morphological details are critical-remains uncertain. The American College of Radiology (ACR) TI-RADS system provides a structured framework for benchmarking artificial intelligence. This study evaluates GPT-5.0's ability to interpret thyroid US images according to TI-RADS criteria and contextualizes its performance relative to expert radiologist assessment, using FNA cytology as the reference standard. Methods: This retrospective study included 100 patients (mean age 49.8 ± 12.6 years; 72 women) with cytology-confirmed diagnoses: Bethesda II (benign) or Bethesda V-VI (malignant). Each nodule had longitudinal and transverse US images acquired with high-frequency linear probes. A board-certified radiologist (>10 years' experience) and GPT-5.0 independently assessed TI-RADS features (composition, echogenicity, shape, margin, echogenic foci) and assigned final categories. Agreement was analyzed using Cohen's κ, and diagnostic performance was calculated using TR4-TR5 as positive for malignancy. Results: Agreement was substantial for composition (κ = 0.62), shape (κ = 0.70), and margin (κ = 0.68); moderate for echogenicity (κ = 0.48); and poor for echogenic foci (κ = 0.12). GPT-5.0 demonstrated a systematic, risk-averse tendency to up-classify nodules, leading to increased TR4-TR5 assignments. Overall, the TI-RADS agreement was 58% (κ = 0.31). The radiologist showed superior diagnostic performance (sensitivity 89%, specificity 85%) compared with GPT-5.0 (sensitivity 67%, specificity 49%), largely driven by false-positive TR4 classifications among benign nodules. Conclusions: GPT-5.0 recognizes several high-level TI-RADS features but struggles with microcalcifications and tends to overestimate malignancy risk within a risk-stratification framework, limiting its standalone clinical use. Ultrasound-specific training and domain adaptation may enable meaningful adjunctive roles in thyroid nodule assessment.
Related Concept Videos
Comparing the Survival Analysis of Two or More Groups
The Thyroid Gland
The follicles have a central cavity lined by simple cuboidal to squamous epithelial cells called follicular cells. These cells produce the glycoprotein...
Interpreting R Charts
An R chart plots the range of subsets of measurements collected from a process. Each point on the chart represents the range—defined as the difference between the maximum and minimum...
Interpreting Run Charts
Functions of Thyroid Hormones
TH is indispensable for the normal development and maturation of the skeletal, muscular, and nervous systems during fetal and childhood growth. It facilitates bone mineral turnover and regulates protein synthesis in developing tissues, contributing significantly to overall growth and...
Interpretation of Confidence Intervals
Confidence intervals have confidence coefficients that are crucial for their interpretation. The most common confidence coefficients are 0.90, 0.95, and 0.99, which can be written as percentages–90%, 95%, and 99%, respectively.
Suppose a person calculates a confidence interval with a confidence coefficient of 0.95. In that case, they can...

