Related Experiment Video
Updated: Mar 12, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Can large language models be trusted? Reliability and readability of responses to perinatal depression FAQs
Jingyu Huang1, Hua Yu2, Junjian Chen3
1Faculty of Health Sciences, University of Macau, Taipa, China.
Large language models (LLMs) show promise in providing reliable information on perinatal depression, but their readability often exceeds public health literacy levels. Further improvements are needed for equitable health communication.
Area of Science:
- Artificial Intelligence in Healthcare
- Medical Informatics
- Public Health Communication
Background:
- Large language models (LLMs) are increasingly utilized in health education, raising concerns about the reliability and readability of AI-generated content.
- Perinatal depression is a significant public health concern, necessitating accessible and accurate information for affected individuals.
Purpose of the Study:
- To evaluate the reliability and readability of responses from five leading LLMs to common questions about perinatal depression.
- To assess whether the readability of AI-generated content aligns with public health literacy standards.
Main Methods:
- Twenty-seven frequently asked questions on perinatal depression were posed to ChatGPT-5, Gemini-2.5, Microsoft Copilot, Grok4, and DeepSeek.
- Responses were independently assessed by two obstetricians using validated instruments (DISCERN, EQIP, JAMA, GQS, HONCODE) for reliability and six indices for readability.
- Inter-rater agreement was quantified using the interclass correlation coefficient (ICC).
Main Results:
- High inter-rater agreement (ICC 0.729–0.847) was observed. Significant differences in reliability scores (DISCERN, EQIP, HONCODE) were found among models (p < 0.001).
- Grok4, DeepSeek, and Copilot showed distinct strengths in specific quality constructs. However, all models exceeded the recommended sixth-grade reading level.
- Readability scores were consistently below benchmarks, indicating potential challenges for individuals with lower health literacy. Most models provided empathetic content but lacked full clinical safety adherence.
Conclusions:
- LLMs demonstrate moderate to high reliability for perinatal depression information, positioning them as potential supplementary resources.
- Readability limitations necessitate improvements to ensure accessibility for all health literacy levels.
- Enhancements in readability, source attribution, and ethical transparency are crucial for maximizing public benefit and achieving equitable health communication.
More Related Videos
06:39Using a Murine Model of Psychosocial Stress in Pregnancy as a Translationally Relevant Paradigm for Psychiatric Disorders in Mothers and Infants
Published on: June 13, 2021
07:31Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Related Concept Videos
Diagnostic and Statistical Manual of Mental Disorders (DSM)
Regression Toward the Mean
Depression: Overview
Long-term Depression
Calcium Ion Concentration Mechanism
If over...
Depressive Disorders: Etiology
Biological Factors in Depression
Biological predispositions significantly influence the risk of developing depressive disorders. Genetic studies highlight the role of variations in the serotonin transporter...