Related Experiment Video
Updated: Aug 6, 2026

13:12
Translational Brain Mapping at the University of Rochester Medical Center: Preserving the Mind Through Personalized Brain Mapping
Published on: August 12, 2019
Patient and physician perspectives on large language model generated responses about brain aneurysm
Joon Hyeok Choi1, Marcella Ruppert-Gomez1, Liam M Shanahan1,2
1Department of Neurosurgery, Boston Children's Hospital, Harvard Medical School, 300 Longwood Ave, Boston, MA, 02115, USA.
Summary
Large language models (LLMs) show variable performance in neurosurgery, with community members finding responses helpful but physicians noting inconsistencies with medical guidelines. Further oversight is needed despite AI advancements.
Area of Science:
- Neurosurgery
- Artificial Intelligence
- Medical Informatics
Background:
- Large language models (LLMs) are increasingly utilized in medicine, including neurosurgery.
- Understanding LLM effectiveness and implications for clinicians and patients is crucial due to their general training.
Purpose of the Study:
- To compare the effectiveness of ChatGPT 4o and Gemini 1.5 Flash in responding to frequently asked questions about brain aneurysms.
- To evaluate community and physician feedback on LLM responses in neurosurgery.
Main Methods:
- External surveys distributed via the Brain Aneurysm Foundation for patients and families.
- Internal surveys completed by neurosurgery physicians at Boston Children's Hospital.
Main Results:
- Community surveys indicated differences in clarity, alternative procedure discussion, and risk discussion between ChatGPT and Gemini.
- Physician surveys revealed significant differences in accuracy, consistency with medical guidelines, omission of key details, and clinical relevance.
- Community participants rated LLM responses more favorably than physicians did.
Conclusions:
- LLM performance in neurosurgery varies, aligning with prior research.
- Physician feedback underscores the necessity of human oversight for LLM-generated medical information.
- A discrepancy exists between community and physician perceptions of LLM response quality.
