Related Experiment Video
Updated: Aug 5, 2026

A Pre-Clinical Model of Synovitis Using Ex vivo Human Synovial Tissue with Preserved Function and Architecture
Published on: March 20, 2026
ChatGPT and Other Large Language Models in Inflammatory Arthritis: A Systematic Review Across Clinical Tasks
Yosef Adiniaev1, Mahmud Omar2, Tohar M Timor3
1Y. Adiniaev, Faculty of Medicine, University of Debrecen, Debrecen, Hungary; BRIDGE GenAI Lab, Boston, Massachusetts, USA.
Objective:
Large language models (LLMs) are increasingly evaluated for rheumatology tasks, but their performance in inflammatory arthritis (IA) remains unclear. We systematically reviewed LLM performance across clinical tasks in IA.
Methods:
We conducted a systematic review (PROSPERO: CRD420261359100), searching PubMed, Scopus, and PubMed Central from January 2022 to April 2026 for studies evaluating LLM performance on clinical tasks in IA. Two reviewers (YA, AG) screened 113 records.
Results:
Eighteen studies covered rheumatoid arthritis (n = 3), axial spondyloarthritis (n = 7), psoriatic arthritis (n = 2), gout (n = 1), juvenile idiopathic arthritis (n = 1), and multiple diseases (n = 4). Most diseases and tasks were represented by only 1 to a few studies, and the evidence base remains early stage and uneven across conditions. Over 20 distinct LLMs were evaluated, including ChatGPT-3.5 to ChatGPT-4o, Gemini 2.0, DeepSeek-R1/V3, Claude, and Perplexity; ChatGPT/GPT variants were the most frequently tested models (16/18 studies), so the current evidence base is predominantly ChatGPT/GPT-based. Findings spanned patient education (n = 11), guideline adherence (n = 6), clinical reasoning (n = 3), and other applications (n = 1). All readability assessments exceeded the recommended thresholds. Guideline concordance ranged from 48% to 96%. Accuracy was lower for case-based clinical scenarios (4.24/6) than for frequently asked questions and guideline-based questions (5.32-5.36/6; P = 0.04). When compared with real clinical data, agreement was poor (Cohen and Fleiss κ ≈ 0).
Conclusion:
LLMs may support patient education, factual medication queries, and structured guideline questions when used under clinician review, but should not be used for case-based reasoning, treatment selection, or autonomous clinical decisions. None of the 18 included studies evaluated retrieval-augmented or agent-based systems, and none prospectively validated LLMs in clinical workflows. Safe integration in rheumatology will require purpose-built, knowledge-grounded systems and prospective evaluation before routine clinical use.
