Related Experiment Video
Updated: May 16, 2026

12:55
Multimodal Protocol for Assessing Metacognition and Self-Regulation in Adults with Learning Difficulties
Published on: September 27, 2020
Assessing large language model responses to pediatric depression FAQs: a cross-sectional study on readability,
RongQi Jiao1, MingZhu Chen1, Jing Zhang1
1Children's Hospital of Nanjing Medical University, Nanjing, China.
Frontiers in Psychiatry
|May 15, 2026
Summary
Large language models (LLMs) offer varying quality for pediatric depression information. DeepSeek 3.1V is most readable, Copilot-5 most accurate, and ChatGPT-5 most complete, but all AI chatbots need human oversight for mental health guidance.
Area of Science:
- Artificial Intelligence in Mental Health
- Pediatric Psychiatry
- Digital Health Information Quality
Background:
- Pediatric depression symptoms vary by age, complicating diagnosis and treatment.
- Parents and adolescents increasingly seek mental health information online, including from AI.
- The quality of online information (readability, accuracy, completeness) is crucial.
Purpose of the Study:
- To compare the suitability of three contemporary large language models (LLMs) as informational tools for pediatric depression.
- To assess the readability, factual accuracy, and completeness of LLM responses to frequently asked questions about pediatric depression.
Main Methods:
- A cross-sectional study analyzed responses from ChatGPT-5, Microsoft Copilot GPT-5, and DeepSeek 3.1V to 15 standardized questions on pediatric depression.
- Readability was assessed using seven indices, accuracy and completeness were scored (0-6 scale), and sentiment analysis was performed.
- Statistical analysis included one-way ANOVA with Tukey post hoc tests.
Main Results:
- Readability varied: DeepSeek 3.1V (Flesch Reading Ease 54-55, Grade Level ~9.5) was easiest to comprehend.
- ChatGPT-5 had intermediate readability (Grade Level ~10.5), while Copilot-5 had the lowest (Grade Level ~10.8).
- Copilot-5 demonstrated the highest accuracy, ChatGPT-5 offered the greatest completeness, and DeepSeek 3.1V excelled in linguistic accessibility.
Conclusions:
- LLMs present diverse capabilities in providing pediatric depression information.
- DeepSeek 3.1V offers better readability, Copilot-5 stronger accuracy, and ChatGPT-5 more comprehensive content.
- AI chatbot information on pediatric mental health requires careful human review and understanding before application.
