Related Experiment Video
Updated: Jun 15, 2025

Author Spotlight: Self-Assessment Protocol for Predicting Psoriatic Arthritis in Psoriasis Patients
Published on: March 1, 2024
Reliability of a generative artificial intelligence tool for pediatric familial Mediterranean fever: insights from a
Saverio La Bella1,2,3, Marina Attanasi4, Annamaria Porreca5
1Department of Pediatrics, "G. D'Annunzio" University of Chieti-Pescara, Chieti, Italy. saveriolabella@outlook.it.
Background:
Artificial intelligence (AI) has become a popular tool for clinical and research use in the medical field. The aim of this study was to evaluate the accuracy and reliability of a generative AI tool on pediatric familial Mediterranean fever (FMF).
Methods:
Fifteen questions repeated thrice on pediatric FMF were prompted to the popular generative AI tool Microsoft Copilot with Chat-GPT 4.0. Nine pediatric rheumatology experts rated response accuracy with a blinded mechanism using a Likert-like scale with values from 1 to 5.
Results:
Median values for overall responses at the initial assessment ranged from 2.00 to 5.00. During the second assessment, median values spanned from 2.00 to 4.00, while for the third assessment, they ranged from 3.00 to 4.00. Intra-rater variability showed poor to moderate agreement (intraclass correlation coefficient range: -0.151 to 0.534). A diminishing level of agreement among experts over time was documented, as highlighted by Krippendorff's alpha coefficient values, ranging from 0.136 (at the first response) to 0.132 (at the second response) to 0.089 (at the third response). Lastly, experts displayed varying levels of trust in AI pre- and post-survey.
Conclusions:
AI has promising implications in pediatric rheumatology, including early diagnosis and management optimization, but challenges persist due to uncertain information reliability and the lack of expert validation. Our survey revealed considerable inaccuracies and incompleteness in AI-generated responses regarding FMF, with poor intra- and extra-rater reliability. Human validation remains crucial in managing AI-generated medical information.
Insights
Generative AI tools show potential in pediatric rheumatology but struggle with accuracy and reliability for conditions like familial Mediterranean fever (FMF). Expert validation is crucial for AI-generated medical information.
Area of Science:
- Pediatric Rheumatology
- Medical Artificial Intelligence
Background:
- Artificial intelligence (AI) is increasingly adopted in clinical and research settings.
- Evaluating AI accuracy in pediatric rheumatology is essential for safe implementation.
Purpose of the Study:
- To assess the accuracy and reliability of a generative AI tool (Microsoft Copilot with Chat-GPT 4.0) for pediatric familial Mediterranean fever (FMF).
Main Methods:
- Fifteen pediatric FMF questions were posed to the AI tool three times.
- Nine pediatric rheumatology experts evaluated response accuracy using a 1-5 Likert-like scale in a blinded manner.
Main Results:
- AI responses showed variable accuracy, with median scores ranging from 2.00 to 5.00 across assessments.
- Poor to moderate intra-rater agreement was observed (ICC: -0.151 to 0.534).
- Expert agreement decreased over time (Krippendorff's alpha: 0.136 to 0.089), and trust in AI varied.
Conclusions:
- While AI offers potential benefits in pediatric rheumatology, significant challenges in information reliability and expert validation exist.
- AI-generated FMF information exhibited inaccuracies and incompleteness, with poor reliability.
- Human expert validation is indispensable for managing AI-derived medical data.

