Related Experiment Video
Updated: May 22, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large language models and child mortality: opportunities and challenges in answering public queries on under-5 causes
Yi Yang1,2,3, Tingxi Zhu4,5, Hongju Chen1,2
1Department of Pediatrics, West China Second University Hospital, Sichuan University, Chengdu, China.
Insights
Large language models (LLMs) show varied performance in child health communication. While generally accurate, their complex language and lack of actionable advice limit public health use.
Area of Science:
- Artificial Intelligence in Healthcare
- Pediatric Public Health Communication
- Digital Health Literacy
Background:
- Reducing under-5 mortality is a global health priority.
- Large language models (LLMs) are increasingly used for public medical information access.
- Evidence on LLMs' performance in child health communication is limited.
Purpose of the Study:
- To evaluate the performance of four leading large language models (LLMs) in responding to public queries on under-5 mortality causes.
- To assess information reliability, accuracy, completeness, comprehensibility, readability, understandability, and actionability of LLM responses.
Main Methods:
- Generated 25 public queries based on top Google Trends search terms for five leading causes of under-5 mortality.
- Collected responses from ChatGPT-4.0, Claude 3.5 Sonnet, Bing AI, and Gemini.
- Evaluated responses using DISCERN, Likert scales, Flesch indices, and PEMAT-P by four pediatricians.
Main Results:
- Significant performance variations observed among LLMs.
- Bing AI scored highest in reliability and overall DISCERN score.
- All models exhibited poor readability (mean FKGL 12.4) and near-zero actionability scores.
Conclusions:
- LLMs provide generally accurate child health information but have limitations in readability and actionability.
- Future LLM development must focus on simplifying language and improving behavioral guidance for effective public health communication.
Background:
Reducing under-5 mortality remains a global health priority. Large language models (LLMs) are increasingly used by the public to access medical information. However, current evidence evaluating LLMs' performance in public-facing child health communication is scarce.
Methods:
We selected the top five search terms related to each of the five leading causes of under-5 mortality (prematurity, pneumonia, birth asphyxia, malaria, and diarrhoea) using Google Trends, generating 25 representative public queries. Responses were collected from four LLMs (ChatGPT-4.0, Claude 3.5 Sonnet, Bing AI, and Gemini) and independently evaluated by four pediatricians. We used the DISCERN instrument for information reliability; 5-point Likert scales for accuracy, completeness, and comprehensibility; Flesch Reading Ease (FRE) and Flesch-Kincaid Grade Level (FKGL) indices for readability; and the Patient Education Materials Assessment Tool for Printable Materials (PEMAT-P) for understandability and actionability. Differences among models were evaluated with Kruskal-Wallis and ANOVA tests, with statistical significance set at p < 0.05.
Results:
We found significant performance variations among the four models across most evaluation metrics. Bing AI achieved the highest total DISCERN score (median 42) and the highest reliability subscore (Section A median 28). Claude consistently underperformed across multiple domains. Notably, readability was poor for all models, with high language complexity (mean FKGL score 12.4). Critically, actionability scores were near zero for all models on the PEMAT-P scale, reflecting a universal lack of clear and practical behavioral guidance.
Conclusion:
While LLMs can generally provide accurate health information, limitations in readability and actionability restrict their practical application in public health communication. Future development should prioritize language simplification and clearer behavioral guidance to enhance their value in public-facing child health communication.
Related Concept Videos
Applications of Life Tables
Life Tables
Population Growth
Steps in Outbreak Investigation
Actuarial Approach
Consider the example of a high-risk surgical procedure with significant early-stage mortality. A two-year clinical study is conducted,...
Truncation in Survival Analysis
Left truncation occurs when individuals who experienced the event of interest before a certain time are not included in the study. This is often due to a "delayed entry" into the study where only those who survive until a certain entry point are observed.