Related Experiment Video
Updated: May 17, 2026

Preliminary Study on Acupuncture Combined with Grain-sized Moxibustion for Treating Rheumatoid Arthritis with Finger Joint Pain
Published on: May 16, 2025
Medical pre-training and fine-tuning improve large-language-model prediction of rheumatoid-arthritis disease activity
Suguru Honda1,2, Katsunori Ikari2,3,4, Mayuko Fujisaki1,2
1Department of Rheumatology, Tokyo Women's Medical University School of Medicine, Tokyo, Japan.
Objective:
To evaluate whether medical-domain pre-training and parameter-efficient fine-tuning improve the ability of locally deployable large language models (LLMs) to predict long-term disease activity and disability in rheumatoid arthritis (RA), and to benchmark their performance against established tabular machine-learning models.
Methods:
We trained on-premises Llama-2 (70B) models with and without medical pre-training (Meditron) and applied QLoRA fine-tuning using structured data from 11 865 patients in the IORRA cohort. Models predicted eight binary outcomes of RA disease activity and disability at 2 years. Performance was compared with logistic regression, random forest, and XGBoost using ROC-AUC, Brier score, calibration plots, and decision-curve analysis.
Results:
Medical pre-training and QLoRA fine-tuning both improved discrimination and calibration, with effects varying by endpoint. Fine-tuned LLMs achieved the highest or comparable ROC-AUCs for most DAS-based outcomes and matched conventional models for remission tasks. In low-prevalence outcomes, LLMs showed more stable calibration and small but consistent net-benefit advantages, while some tabular models exhibited clinically harmful decision curves.
Conclusions:
Endpoint-specific, privacy-preserving LLMs can complement or replace conventional tabular models for RA disease-activity prediction, providing reliable, locally deployable decision support without sacrificing data privacy.
