Related Experiment Video
Updated: Oct 10, 2026

The bm12 Inducible Model of Systemic Lupus Erythematosus (SLE) in C57BL/6 Mice
Published on: November 1, 2015
Performance of Large Language Models in Answering Juvenile Systemic Lupus Erythematosus Questions: A Blinded
Ufuk Furkan Ozdemir1, Serpil Meric Toprak1, Aybuke Gunalp2
1Department of Pediatric Rheumatology, Istanbul Medeniyet University.
Background:
Large language models are increasingly used to obtain medical information, but their performance in juvenile systemic lupus erythematosus remains insufficiently characterized.
Objective:
To compare expert-rated response quality and clinical relevance across 4 artificial intelligence (AI) models using standardized juvenile systemic lupus erythematosus-related questions.
Design:
Cross-sectional expert-based comparative study.
Setting:
Freely accessible consumer-facing web interfaces evaluated on December 30, 2025.
Participants:
Ten pediatric rheumatology experts.
Intervention:
Not applicable.
Main Outcome Measures:
Twenty standardized questions generated 80 AI responses, rated on a 5-point Likert Scale. Model performance was compared using the Friedman test with Kendall W and Bonferroni-corrected post hoc analyses; interrater agreement was assessed using intraclass correlation coefficients (ICCs).
Results:
DeepSeek achieved the highest median rating [5 (1 to 5)], followed by Gemini [4 (2 to 5)], ChatGPT [4 (3 to 5)], and Copilot [4 (3 to 5)]. Overall expert-rated performance differed significantly among models [Friedman χ2(3) = 63.553, p < 0.001; Kendall W = 0.106]. DeepSeek outperformed ChatGPT, Gemini, and Copilot (all p < 0.001), while Gemini performed better than ChatGPT (p = 0.007) and Copilot (p < 0.001). Interrater agreement was highest for DeepSeek (ICC =0 .734) and lowest for Copilot (ICC = 0.043). Trust scores remained unchanged before and after evaluation [4 (3 to 5) vs. 4 (3 to 5), p=0.414].
Limitation:
Single-time-point, single-response testing did not assess within-model reproducibility, and the question set was not formally validated.
Conclusion:
AI models differed in expert-rated performance and interrater agreement; outputs should be interpreted cautiously and with specialist oversight.