Reliability of a generative artificial intelligence tool for pediatric familial Mediterranean fever: insights from a

Saverio La Bella1,2,3, Marina Attanasi4, Annamaria Porreca5

  • 1Department of Pediatrics, "G. D'Annunzio" University of Chieti-Pescara, Chieti, Italy. saveriolabella@outlook.it.

Abstract

Insights

Generative AI tools show potential in pediatric rheumatology but struggle with accuracy and reliability for conditions like familial Mediterranean fever (FMF). Expert validation is crucial for AI-generated medical information.

Area of Science:

  • Pediatric Rheumatology
  • Medical Artificial Intelligence

Background:

  • Artificial intelligence (AI) is increasingly adopted in clinical and research settings.
  • Evaluating AI accuracy in pediatric rheumatology is essential for safe implementation.

Purpose of the Study:

  • To assess the accuracy and reliability of a generative AI tool (Microsoft Copilot with Chat-GPT 4.0) for pediatric familial Mediterranean fever (FMF).

Main Methods:

  • Fifteen pediatric FMF questions were posed to the AI tool three times.
  • Nine pediatric rheumatology experts evaluated response accuracy using a 1-5 Likert-like scale in a blinded manner.

Main Results:

  • AI responses showed variable accuracy, with median scores ranging from 2.00 to 5.00 across assessments.
  • Poor to moderate intra-rater agreement was observed (ICC: -0.151 to 0.534).
  • Expert agreement decreased over time (Krippendorff's alpha: 0.136 to 0.089), and trust in AI varied.

Conclusions:

  • While AI offers potential benefits in pediatric rheumatology, significant challenges in information reliability and expert validation exist.
  • AI-generated FMF information exhibited inaccuracies and incompleteness, with poor reliability.
  • Human expert validation is indispensable for managing AI-derived medical data.

Related Concept Videos