Related Experiment Video
Updated: May 5, 2026

Author Spotlight: Self-Assessment Protocol for Predicting Psoriatic Arthritis in Psoriasis Patients
Published on: March 1, 2024
Evaluating the Accuracy, Readability, and Consistency of Artificial Intelligence Models in Patient Education for
1Department of Engineering Medicine, Texas A&M University, Houston, USA.
Introduction And Aim:
Artificial intelligence (AI) chatbots are increasingly used by patients to obtain medical information before seeking clinical care; however, the accuracy, readability, and consistency of AI-generated information in rheumatology remain uncertain. This study aimed to assess the readability, accuracy, and consistency of large language models (LLMs) generated patient education content for rheumatoid arthritis (RA), osteoarthritis (OA), and psoriatic arthritis (PsA).
Methods:
From August 18, 2025, to August 24, 2025, three standardized patient-facing questions per disease were submitted daily to three LLMs, specifically ChatGPT (San Francisco, CA: OpenAI), Google Gemini (Mountain View, CA: Google LLC), and OpenEvidence (Cambridge, MA: OpenEvidence Inc.), with histories cleared between submissions. Readability (Hemingway grade level), word count, accuracy (on a 1-5 scale), and day-to-day consistency (Jaccard similarity) were measured. Responses from each model's most consistent day were accuracy-rated by rheumatology fellows and attendings.
Results:
A total of 189 responses were collected. No model consistently met the American Medical Association (AMA)/National Institutes of Health (NIH)-recommended reading levels for sixth through eighth grade. OpenEvidence produced the most technical content (≥17th grade, indicating postbaccalaureate readability), while ChatGPT and Gemini averaged 11.9-12.0 grade. Gemini generated the longest responses (>500 words). ChatGPT showed the highest day-to-day stability (range: <0.07), Gemini moderate variability, and OpenEvidence both the widest range (0.24) and the highest average similarity (0.383). Accuracy ratings varied as follows: OpenEvidence generally scored higher for RA and PsA, while ChatGPT and Gemini were similar across diseases. OA responses showed minimal difference. Best and worst responses mirrored these trends.
Conclusions:
Current LLMs generate rheumatology information above the recommended reading levels. ChatGPT was the most consistent, Gemini the most detailed, and OpenEvidence the most technical. Persistent barriers to readability highlight the need for health-literacy-optimized AI communication.
