Related Experiment Video
Updated: Sep 5, 2026

Synchronous Triplanar Reconstruction Integrated with Color Doppler Mapping for Precise and Rapid Localization of Thyroid Lesions
Published on: February 9, 2024
Patient-perceived quality of artificial intelligence responses in Hashimoto's thyroiditis
Rıfat Furkan Aydın1, Özge Baş Aksu2, Sevim Güllü2
1Ankara University, School of Medicine, Department of Internal Medicine - Ankara, Türkiye.
Objective:
Artificial intelligence-driven conversational models are increasingly used for patient education, yet whether information patients find clear and satisfying also meets expert accuracy standards remains unclear. The aim of this study was to compare patient and physician evaluations of ChatGPT-5 and DeepSeek V3.1 responses to questions on Hashimoto's thyroiditis.
Methods:
In a cross-sectional, double-blind, within-participant study, twenty standardized, endocrinologist-developed questions across six domains were submitted to both models in October 2025. Seventeen adults with confirmed Hashimoto's thyroiditis (mean age 47.6±13.5 years; 88.2% female; 64.7% bachelor's degree or higher) rated each blinded response for clarity and satisfaction on 10-point Likert scales and selected a preferred response. Two endocrinologists also rated responses blindly; with only two physicians, their ratings were summarized descriptively. Paired t-tests with Holm correction compared ratings, and preference was analyzed with mixed-effects logistic regression.
Results:
Patients rated ChatGPT-5 higher than DeepSeek V3.1 for clarity (9.15±0.70 vs. 8.41±1.14; Holm-adjusted p=0.005; Cohen's d_z=0.87) and satisfaction (8.93±0.94 vs. 8.19±1.37; Holm-adjusted p=0.005; d_z=0.85), and preferred ChatGPT-5 in 68.8% of selections (odds ratio 2.20; 95%CI 1.75-2.77; p<0.001). In exploratory analyses, this advantage was concentrated among higher-education participants and was undetectable in the small lower-education subgroup (n=6). Owing to the small physician sample (n=2) and low inter-rater agreement (κ=0.12; accuracy intraclass correlation coefficient=0.48), physician accuracy ratings (5.90 and 5.70/10) are exploratory and do not support a robust patient-physician comparison.
Conclusion:
Patients preferred ChatGPT-5 for clarity and satisfaction, though this advantage may not extend to lower-health-literacy populations; high patient satisfaction should not be taken as a proxy for clinical accuracy. Physician-supervised deployment and language-specific validation are recommended for responsible integration of artificial intelligence into patient education.
Related Concept Videos
Graves' Disease I: Introduction
Hyperthyroidism II: Pathophysiology