Related Experiment Video
Updated: Aug 5, 2026

Preliminary Study on Acupuncture Combined with Grain-sized Moxibustion for Treating Rheumatoid Arthritis with Finger Joint Pain
Published on: May 16, 2025
Large language model-based extraction of rheumatoid arthritis clinical disease activity index from clinical notes
Reid Weisberg1, Ruoqi Yang2, Iram Kamdar2
1Division of Rheumatology, Department of Medicine, UT Southwestern Medical Center, Dallas, TX 10032, United States.
Objectives:
We extracted a validated disease activity measure in rheumatoid arthritis (RA), the Clinical Disease Activity Index (CDAI), from a large tertiary academic medical center electronic health record (EHR) using an automated large language model (LLM)-based approach without requiring model pretraining.
Materials And Methods:
The New York Presbyterian/Columbia University Medical Center Clinical Data Warehouse contains EHR data for over 4.5 million patients. RA patients were identified using International Classification of Disease-9 (ICD-9) and ICD-10 codes. Expert-curated CDAI keywords were extracted from unstructured notes using an automated natural language processing (NLP) pipeline leveraging GPT-4o API, a HIPAA-compliant, institutionally approved LLM platform. Performance was evaluated against expert chart review.
Results:
Among 2756 RA patients with notes, 1038 (37.7%) were seropositive, 796 (28.9%) were seronegative, and 922 (33.4%) had unknown serostatus. Clinical Disease Activity Index and its components were extracted in 15.4% (160/1038) of seropositive patients indicating remission or low disease activity. Clinical Disease Activity Index documentation was more frequent among patients with multiple notes and among faculty, with high extraction accuracy (precision/recall/F1 = 0.97).
Discussion:
This represents the first attempt to employ a zero-shot, ChatGPT-powered LLM platform to extract RA disease activity measures from real-world EHR data. Although a low prevalence of documentation was noted, important distinctions were observed when patients were subgrouped by serostatus, level of training, and number of visits.
Conclusion:
An LLM-based pipeline accurately extracted CDAI from a single large academic EHR, revealing infrequent real-world documentation.

