Related Experiment Video
Updated: May 5, 2026

14:32
Using Visual and Narrative Methods to Achieve Fair Process in Clinical Care
Published on: February 16, 2011
24.3K
Artificial Intelligence Chatbot Responses to Patient Queries on Traumatic Brain Injury: An Expert Assessment of
Patrick Schuss1, Andreas S Gonschorek2, Michael Kämper3
1Department of Neurosurgery, BG Klinikum Unfallkrankenhaus Berlin, Berlin, Germany.
Journal of Neurotrauma
|December 3, 2025
Summary
Artificial intelligence chatbots like ChatGPT show promise in answering medical questions about traumatic brain injury (TBI), but require further development for empathy and context awareness before widespread clinical use.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Neurology
Background:
- Growing use of AI chatbots for medical inquiries necessitates evaluation of their accuracy and reliability.
- Patient education and information access are critical aspects of healthcare delivery.
Purpose of the Study:
- To systematically evaluate the performance of ChatGPT, Google Gemini, and Microsoft CoPilot in answering patient-oriented questions about traumatic brain injury (TBI).
- To compare the accuracy, reliability, and trustworthiness of AI chatbot responses against expert-prepared answers and patient perceptions.
Main Methods:
- A standardized set of TBI-related questions across eight subtopics was posed to three AI chatbots using unified prompts.
- Chatbot responses were evaluated against expert-generated reference answers and assessed by TBI rehabilitation patients.
- Performance was measured using a modified scoring framework across five quality dimensions, with statistical analysis including ANOVA and logistic regression.
Main Results:
- Significant differences were observed in chatbot performance across quality dimensions.
- ChatGPT demonstrated higher scores in reliability, responsiveness, and perceived trustworthiness compared to Gemini and CoPilot (p < 0.05).
- ChatGPT responses were significantly more likely to be considered adequate substitutes for expert advice (p < 0.0001, OR = 4.3).
Conclusions:
- AI-driven chatbots exhibit varying capabilities in delivering high-quality medical information, with notable differences in reliability and responsiveness.
- While ChatGPT shows an advantage in structured information delivery for TBI queries, improvements in contextual understanding and empathy are crucial for clinical integration.
- Further research is needed to refine AI chatbot performance for safe and effective use in patient education and medical information dissemination.

