Related Experiment Video
Updated: Apr 9, 2026

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
Published on: April 21, 2023
Evaluating Chain-of-Thought reasoning in large language models for thyroid ultrasound interpretation: a
Yu-Tong Zhang1, Si-Yi Wu1, Dong Zhang2
1The Department of Ultrasound, The Second Affiliated Hospital of Xi'an Jiaotong University, Xi'an, Shaanxi, China.
Objective:
To assess whether reasoning-capable large language models (LLMs) can accurately interpret both qualitative and quantitatively encoded ultrasound features of thyroid nodules within the ACR-TIRADS framework and improve diagnostic reliability.
Methods:
This retrospective study analyzed thyroid nodules with both radiologist-labeled qualitative ultrasound features and quantitatively encoded descriptors generated through standardized numerical modeling. Both formats were converted into structured prompts and input separately into four CoT-enabled LLMs (ChatGPT-O3, Grok-3, DeepSeek-R1, Gemini-2.5 Pro), each performing three reasoning rounds per task. Diagnostic performance was evaluated by accuracy and reproducibility, and two types of inconsistencies-cross-threshold and cross-modal conflicts-were quantified. Reasoning authenticity and conciseness were independently assessed by radiologists of varying experience. Sankey diagrams were used to summarize ACR-TIRADS category transitions.
Results:
ChatGPT-O3, Gemini-2.5 Pro, and Grok-3 showed strong ACR-TIRADS accuracy (91, 96, 96%), outperforming DeepSeek-R1 (79%). Grok-3 was highest in score-based accuracy (96%); DeepSeek-R1 lowest (52%). Reproducibility for categorization was Grok-3 93%, Gemini-2.5 Pro 90%, ChatGPT-O3 88%, vs. DeepSeek-R1 67%. For scoring reproducibility, Grok-3 (93%), ChatGPT-O3 (90%), and Gemini-2.5 Pro (79%) exceeded DeepSeek-R1 (18%). Physicians rated Grok-3 and Gemini-2.5 Pro highest in reasoning authenticity, while ChatGPT-O3 was most concise (mean 144 words). For quantitative tasks, Gemini-2.5 Pro (78%) and DeepSeek-R1 (74%) were most accurate; Grok-3 lowest (64%). Reproducibility was highest for Gemini-2.5 Pro (84%) and DeepSeek-R1 (86%). Across models, the proportion of nodules exhibiting cross-threshold discrepancies ranged from 3 to 17%, with Grok-3 lowest and DeepSeek-R1 highest. Cross-modal conflicts were more frequent, ranging from 27 to 36% across the four LLMs.
Conclusion:
Grok-3 excelled in qualitative tasks, while Gemini-2.5 Pro and DeepSeek-R1 showed strengths in quantitative analysis. CoT-enabled LLMs offered interpretable reasoning with promise for clinical decision support.
Related Concept Videos
Inductive Reasoning
Inductive reasoning is common in descriptive science. A life scientist makes observations and records them. This data can be qualitative or...
Reasoning
Inductive reasoning involves deriving generalizations from specific observations. This type of reasoning helps form beliefs about the world. For example,...
Deductive Reasoning
For example, a researcher can deduce specific predictions...
Language and Cognition
Critical Thinking II
Synthesis and Regulation of Thyroid Hormones
Upon reaching the thyroid gland, TSH stimulates the follicular cells' active uptake of iodide ions from the blood. The ions diffuse to the apical surface of the cells and are oxidized to iodine. The...

