Related Experiment Video
Updated: Jun 14, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
How do large language models answer ADHD-related questions? A comparative study of ChatGPT, Gemini, and DeepSeek
Berrin Bilgiç1, Serkan Turan2, Sibelnur Avcil3
1Department of Child and Adolescent Psychiatry, Faculty of Medicine, Adnan Menderes University, Aydın, Türkiye. berrinbilgic@adu.edu.tr.
Objective:
Large language models (LLMs) are increasingly used by patients and caregivers as sources of health information. However, their performance in addressing attention-deficit/hyperactivity disorder (ADHD)-related questions has not been systematically compared. This study aimed to evaluate and compare the accuracy, reproducibility, quality, usefulness, and reliability of responses generated by ChatGPT (GPT-4o), Gemini, and DeepSeek R1.
Methods:
In this cross-sectional comparative study, 22 commonly asked ADHD-related questions identified from publicly available digital sources were categorized into four domains: basic knowledge, diagnosis and assessment, treatment and medication, and long-term outcomes. Each question was presented to all three models using the same standardized prompts in separate chat sessions. The generated responses were independently evaluated by two specialists in child and adolescent psychiatry. Reproducibility was examined by repeating the same queries on different days. Descriptive statistics and non-parametric repeated-measures analyses were used to compare model performance.
Results:
All models showed high overall accuracy, with mean scores of 91% for ChatGPT (GPT-4o), 89% for Gemini, and 87% for DeepSeek R1. Reproducibility followed a similar pattern (89%, 86%, and 84%, respectively). Gemini and DeepSeek performed relatively better in basic knowledge and diagnostic domains, whereas ChatGPT (GPT-4o) showed stronger performance in treatment and long-term outcome-related questions. Significant differences were observed in quality, usefulness, and reliability across models, with ChatGPT (GPT-4o) achieving the highest overall expert-rated scores.
Conclusion:
Although large language models generally provided accurate responses to ADHD-related questions, notable differences were observed in the depth, clarity, and clinical usefulness of the information across models. These systems may serve as supportive sources of information for patients and caregivers; however, their responses should be interpreted with caution and should not replace professional clinical evaluation or medical advice.