Related Experiment Video
Updated: Jul 11, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Assessment of CDASI scoring by a multimodal large language model: a comparative study with expert assessors
Marco Fornaro1, Vincenzo Venerito1, Swapnasha Panigrahi2
1Unit of Rheumatology, Department of Precision and Regenerative Medicine, University of Bari, Area Jonica (DiMePRe-J), Bari, Italy.
Rheumatology International
|July 8, 2026
Summary
Large language models show promise in assessing dermatomyositis (DM) activity using the CDASI tool. While efficient, LLMs need physician oversight for accurate damage assessment in DM.
Area of Science:
- Dermatology
- Artificial Intelligence
- Rheumatology
Background:
- Cutaneous disease activity significantly impacts dermatomyositis (DM) patients.
- Accurate assessment of DM disease activity is crucial for treatment decisions.
- Automated tools may enhance the efficiency and consistency of disease scoring.
Purpose of the Study:
- To evaluate the performance of a multimodal large language model (LLM), Claude v3.5 Sonnet, in scoring the Cutaneous Dermatomyositis Disease Area and Severity Index (CDASI).
- To compare the LLM's scoring accuracy against expert rheumatologists for both disease activity and damage in DM.
- To assess the time efficiency of LLM-based CDASI scoring.
Main Methods:
- Retrospective analysis of 30 published DM cases with clinical images.
- Independent scoring of CDASI by two expert rheumatologists.
- LLM assessment using structured prompting based on CDASI definitions.
- Intraclass correlation coefficients (ICCs) used to evaluate agreement.
Main Results:
- Claude v3.5 Sonnet demonstrated good agreement for CDASI activity (ICC 0.71) but lower reliability for damage assessment (ICC 0.41).
- Moderate agreement was observed for activity domains like erythema and scaling; lower concordance for chronic damage features like poikiloderma.
- Hand assessments, particularly periungual changes (ICC 1.0) and global hand scores (ICC 0.95), showed strong performance.
- LLM evaluation time was reduced by approximately 92% compared to clinicians (42s vs. 8.4 min per case).
Conclusions:
- Multimodal LLMs show potential for automated CDASI activity assessment in DM, with promising agreement with expert raters.
- LLMs significantly improve scoring efficiency, reducing evaluation time.
- Lower reliability in damage assessment highlights the continued necessity of physician oversight in DM management.
