Related Experiment Video
Updated: Jul 16, 2026

IntelliSleepScorer, a Software Package with a Graphic User Interface for Mice Automated Sleep Stage Scoring
Published on: November 8, 2024
How Correct is AI for Infant Safe Sleep Advice? Evaluating Accuracy of ChatGPT, Gemini, and Claude Against AAP
Evin Rothschild1, Jack Christian1, Celine Arar1
1Johns Hopkins School of Medicine, 733 N Broadway, Baltimore, Maryland.
Insights
Large language models (LLMs) provide inconsistent infant safe sleep advice, though they are empathetic. Gemini demonstrated higher accuracy than ChatGPT, but pediatric oversight is crucial for evidence-based online information.
Area of Science:
- Pediatric Sleep Medicine
- Artificial Intelligence in Healthcare
- Health Information Technology
Background:
- Caregivers increasingly use online resources for infant safe sleep guidance.
- Pediatricians promote evidence-based safe sleep practices to reduce infant mortality.
- Previous research has not comprehensively evaluated large language model (LLM) accuracy for infant safe sleep information.
Purpose of the Study:
- To assess the accuracy of LLM responses to caregiver questions on infant safe sleep.
- To compare LLM-generated advice against the American Academy of Pediatrics (AAP) 2022 recommendations.
- To evaluate LLM performance in terms of completeness, empathy, and response stability.
Main Methods:
- Nine caregiver questions from Reddit were used to query three LLMs (ChatGPT, Gemini, Claude).
- Responses were evaluated for accuracy, completeness, and empathy using a 0-2 scale by three reviewers.
- Readability was assessed using Flesch-Kincaid grade level; response stability was measured across multiple queries.
Main Results:
- Gemini (mean 1.85) showed significantly higher accuracy than ChatGPT (1.30) (p=0.01).
- All models scored high for empathy (mean 2) and had comparable completeness and stability.
- Readability levels ranged from grade 7 to 9, with ChatGPT being the most readable.
Conclusions:
- LLMs provide infant safe sleep advice that is empathetic but inconsistently accurate, often deviating from AAP guidelines.
- Direct questions about guidelines were more accurate than nuanced queries.
- Pediatrician oversight and collaboration with AI developers are vital to ensure families receive safe, evidence-based online information.
Objective:
Since the 1994 "Back to Sleep" campaign, pediatricians have promoted evidence-based infant safe sleep practices to reduce sleep-related infant deaths. However, caregivers increasingly seek guidance online. We sought to determine the accuracy of large language model (LLM) responses to caregiver questions about infant safe sleep, compared with the American Academy of Pediatrics' (AAP) 2022 recommendations.
Design:
Nine caregiver questions adapted from Reddit New Parents forum were mapped to core AAP safe sleep topics. Each was entered into three LLMs: ChatGPT 5, Gemini 2.5 Flash, and Claude Sonnet 4.5, three times within the same day to assess stability. Three reviewers scored responses on a 0-2 scale for accuracy (primary outcome), completeness, and empathy. Stability reflected similarity across repeated responses. Readability was calculated using the Flesch-Kincaid grade level. Mean scores were compared using descriptive statistics, analysis of variance, and post hoc testing.
Results:
Mean accuracy varied significantly across models. Gemini had the highest accuracy score (mean 1.85), followed by Claude (1.44), and ChatGPT (1.30). Gemini was significantly more accurate than ChatGPT (p=0.01). All models scored high in empathy (2). There were no significant differences in completeness and stability between models. ChatGPT had the lowest average readability, with all models' reading levels between grades seven to nine (7.64 vs 9.15 vs 8.82, p=0.01). Direct guideline questions yielded higher accuracy than nuanced questions.
Conclusion:
LLMs offer inconsistently accurate but empathetic infant safe sleep advice with frequent deviations from AAP recommendations. Pediatric oversight and collaboration with technology developers are essential to ensure safe, evidence-based information for families.

