Related Experiment Video
Updated: Jul 1, 2026

Generation of a Mouse Spontaneous Autoimmune Thyroiditis Model
Published on: March 17, 2023
Risk-centered benchmarking of large language models for AI-enabled counseling in chronic autoimmune thyroid eye
Fangqin Fei1, Lu Xie2, Jing Rao2
1Department of Endocrinology, The First People's Hospital of Huzhou, Huzhou Normal University, Huzhou, Zhejiang, China.
Background:
Thyroid eye disease (TED) is a chronic autoimmune inflammatory orbital disease requiring activity assessment, risk stratification, and triage. As patients increasingly consult large language models (LLMs), evidence on their quality and safety for TED counseling remains limited.
Methods:
We conducted a cross-sectional benchmark using a prespecified 35-question Chinese TED counseling bank covering symptom recognition, activity assessment, treatment, daily management, follow-up, and care-seeking. The bank was developed from guideline-/consensus-derived scenarios, expert discussion, and recurrent patient-inquiry themes, then applied under a unified single-turn protocol. Five web-based LLM chatbots were evaluated: Gemini 3 Pro, ChatGPT-5.2, DeepSeek-V3.1, Doubao, and Qwen3-Max. Systems were accessed through official interfaces in Quzhou, China, during 27-29 December 2025, with identifiers recorded. Automated text analysis extracted output features, and response time was measured. Two blinded expert raters assessed alignment with a guideline-/consensus-informed reference standard using 5-point Likert scales for Accuracy, Logic, Coherence, Safety, and Content Accessibility. Between-model comparisons used repeated-measures methods, and correlations used Spearman analysis.
Results:
Response time differed significantly across models (Friedman χ2 = 94.79, P < 0.001), with Gemini 3 Pro fastest and Doubao slowest. Output characteristics varied: Doubao generated the longest responses, ChatGPT-5.2 the shortest, and Qwen3-Max the most table-formatted outputs. Significant between-model differences were found for Accuracy (χ2 = 16.64, P = 0.002), Logic (χ2 = 20.76, P < 0.001), Coherence (χ2 = 15.54, P = 0.004), and Content Accessibility (χ2 = 23.51, P < 0.001), but not Safety (χ2 = 1.03, P = 0.905). Response time correlated moderately with output length (words ρ = 0.53; characters ρ = 0.51). Content Accessibility correlated weakly with length and table use (tables ρ = 0.29), whereas longer outputs were not consistently associated with higher Accuracy or Logic.
Conclusion:
LLMs show marked heterogeneity in efficiency, output structure, and clinical quality for TED counseling. Longer or slower responses do not necessarily indicate better performance. The absence of between-model Safety differences should not be interpreted as absolute safety or equivalence. Risk-centered, structured outputs emphasizing red-flag symptoms and care-seeking thresholds warrant validation through multi-turn dialogues, repeated sampling, patient/lay-user evaluation, and finer-grained safety endpoints.
Related Concept Videos
Graves' Disease I: Introduction
Synthesis and Regulation of Thyroid Hormones
Upon reaching the thyroid gland, TSH stimulates the follicular cells' active uptake of iodide ions from the blood. The ions diffuse to the apical surface of the cells and are oxidized to iodine. The iodine is then...
Graves Disease II: Pathophysiology
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in situations...
Autoimmune Disorders
Concept and Mechanism of Autoimmune Diseases
The immune system...