Related Experiment Video
Updated: Apr 21, 2026

A Computer-Based Platform for Aiding Clinicians in Eating Disorder Analysis and Diagnosis
Published on: May 10, 2022
A Comparative Evaluation of Three Large Language Models for Parent-Centered Questions About Anorexia Nervosa
Celal Yeşilkaya1, Hande Kırışman Keleş2, Esra Rabia Taşpolat3
1Department of Child and Adolescent Psychiatry, Konya Ereğli State Hospital, Konya, Türkiye.
Large language models (LLMs) offer generally reliable health information for parents regarding anorexia nervosa (AN). However, AI-generated guidance should complement, not replace, professional clinical advice due to potential omissions and inaccuracies.
Area of Science:
- Artificial Intelligence in Healthcare
- Child and Adolescent Mental Health
Background:
- Large language models (LLMs) are increasingly utilized for health information, including child and adolescent mental health.
- Anorexia nervosa (AN) requires early recognition and intervention, making accurate AI-generated information crucial for parents.
- This study assessed the performance of LLMs in answering parent queries about AN.
Purpose of the Study:
- To evaluate the accuracy and reliability of conversational AI systems in responding to parent-oriented questions about anorexia nervosa (AN).
- To compare the performance of ChatGPT (GPT-4o), Google Gemini, and DeepSeek in providing information on AN for parents.
- To identify limitations and areas for improvement in AI-generated health information for AN.
Main Methods:
- Comparative model evaluation of three LLMs: ChatGPT (GPT-4o), Google Gemini, and DeepSeek.
- Twenty representative parent questions about AN were curated and submitted using standardized prompts.
- Responses were anonymized and evaluated by two child and adolescent psychiatrists for quality, usefulness, and reliability, with reproducibility assessed via repeated queries.
Main Results:
- All evaluated LLMs demonstrated high reliability and overall performance.
- ChatGPT showed the highest accuracy (≈92%) and reproducibility (≈90%), followed by Gemini (≈88% accuracy; ≈85% reproducibility) and DeepSeek (≈86% accuracy; ≈83% reproducibility).
- Models exhibited lower accuracy for diagnosis and clinical assessment questions, with common limitations including omitted information, factual inaccuracies (DeepSeek), and generalized guidance (Gemini).
Conclusions:
- LLMs can offer broadly accurate preliminary information for parents seeking AN guidance.
- Limitations such as information omissions and domain-specific variability necessitate caution.
- AI-generated information should serve as a supplementary resource, not a substitute for professional clinical guidance.
More Related Videos
07:56Assessing the Coherence of Parents' Short Narratives Regarding their Child Using the Five-Minute Speech Sample Procedure
Published on: September 19, 2019
06:15Using the Visual World Paradigm to Study Sentence Comprehension in Mandarin-Speaking Children with Autism
Published on: October 3, 2018
Related Concept Videos
Anorexia Nervosa
Symptoms and Physical Effects
Individuals with anorexia nervosa commonly exhibit extreme...
Bulimia Nervosa
Binge Eating Disorders
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...