How Correct is AI for Infant Safe Sleep Advice? Evaluating Accuracy of ChatGPT, Gemini, and Claude Against AAP

Evin Rothschild1, Jack Christian1, Celine Arar1

  • 1Johns Hopkins School of Medicine, 733 N Broadway, Baltimore, Maryland.

Academic Pediatrics
|July 14, 2026
PubMed

Insights

Large language models (LLMs) provide inconsistent infant safe sleep advice, though they are empathetic. Gemini demonstrated higher accuracy than ChatGPT, but pediatric oversight is crucial for evidence-based online information.

Area of Science:

  • Pediatric Sleep Medicine
  • Artificial Intelligence in Healthcare
  • Health Information Technology

Background:

  • Caregivers increasingly use online resources for infant safe sleep guidance.
  • Pediatricians promote evidence-based safe sleep practices to reduce infant mortality.
  • Previous research has not comprehensively evaluated large language model (LLM) accuracy for infant safe sleep information.

Purpose of the Study:

  • To assess the accuracy of LLM responses to caregiver questions on infant safe sleep.
  • To compare LLM-generated advice against the American Academy of Pediatrics (AAP) 2022 recommendations.
  • To evaluate LLM performance in terms of completeness, empathy, and response stability.

Main Methods:

  • Nine caregiver questions from Reddit were used to query three LLMs (ChatGPT, Gemini, Claude).
  • Responses were evaluated for accuracy, completeness, and empathy using a 0-2 scale by three reviewers.
  • Readability was assessed using Flesch-Kincaid grade level; response stability was measured across multiple queries.

Main Results:

  • Gemini (mean 1.85) showed significantly higher accuracy than ChatGPT (1.30) (p=0.01).
  • All models scored high for empathy (mean 2) and had comparable completeness and stability.
  • Readability levels ranged from grade 7 to 9, with ChatGPT being the most readable.

Conclusions:

  • LLMs provide infant safe sleep advice that is empathetic but inconsistently accurate, often deviating from AAP guidelines.
  • Direct questions about guidelines were more accurate than nuanced queries.
  • Pediatrician oversight and collaboration with AI developers are vital to ensure families receive safe, evidence-based online information.
Abstract