Related Experiment Video
Updated: Apr 22, 2026

Author Spotlight: Integrating Ultrasound Imaging with Biochemical Markers for Thyroid Disease Diagnosis
Published on: February 9, 2024
A Comparative Assessment of Large Language Models in Congenital Hypothyroidism: Reliability, Quality and Readability
Ebru Barsal Çetiner1, Berna Singin1
1University of Health Sciences Türkiye, Antalya City Hospital, Clinic of Pediatric Endocrinology, Antalya, Türkiye
Large language models (LLMs) provide moderate-to-high quality information on congenital hypothyroidism (CH), but readability is often too complex for patients. ChatGPT-5.2 performed best, though all models exceeded recommended reading levels.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Patient Education
Background:
- Congenital hypothyroidism (CH) requires accurate and accessible patient information.
- Large language models (LLMs) are increasingly used for health information retrieval.
- Evaluating the performance of LLMs in providing CH information is crucial for patient understanding.
Purpose of the Study:
- To compare the reliability, quality, and readability of responses from leading LLM-based chatbots regarding congenital hypothyroidism (CH).
- To assess the suitability of LLM-generated CH content for patient education.
- To identify the best-performing LLM for CH-related patient queries.
Main Methods:
- Forty clinician-vetted CH frequently asked questions (FAQs) were posed to ChatGPT-4, ChatGPT-5.2, Gemini, and Copilot.
- Reliability was assessed using the modified DISCERN (mDISCERN) instrument.
- Quality was evaluated using the Global Quality Score (GQS), and readability by multiple indices (FRE, FKGL, GFI, CLI, SMOG).
Main Results:
- ChatGPT-5.2, ChatGPT-4, and Gemini achieved higher median mDISCERN and GQS scores than Copilot.
- ChatGPT-5.2 demonstrated superior performance in reliability, quality, and readability metrics.
- All evaluated LLMs produced content exceeding the recommended sixth-grade reading level, indicating suboptimal readability for patient education.
Conclusions:
- LLM-based chatbots can generate moderately to highly reliable and quality information on CH.
- Readability of LLM-generated CH content is a significant limitation for patient comprehension.
- While LLMs can supplement patient information needs, they should not replace professional medical advice and counseling.
Related Concept Videos
Hypothyroidism II: Pathophysiology
Synthesis and Regulation of Thyroid Hormones
Upon reaching the thyroid gland, TSH stimulates the follicular cells' active uptake of iodide ions from the blood. The ions diffuse to the apical surface of the cells and are oxidized to iodine. The...
Hyperthyroidism II: Pathophysiology
Hyperthyroidism I: Introduction
Language and Cognition
Goiter

