Related Experiment Video
Updated: Jul 7, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Utilizing Artificial Intelligence and Chat Generative Pretrained Transformer to Answer Questions About Clinical
Samuel N Blacker1, Mia Kang1, Indranil Chakraborty2
1Department of Anesthesiology, University of North Carolina at Chapel Hill.
Artificial intelligence chatbot ChatGPT demonstrated limited reliability in answering clinical guideline questions. Targeted questions improved accuracy, but ChatGPT is not yet suitable for clinical decision-making.
Area of Science:
- Neuroscience
- Artificial Intelligence
- Clinical Guidelines
Background:
- Clinical guidelines are essential for evidence-based medical practice.
- Artificial intelligence (AI) chatbots like ChatGPT are increasingly used for information retrieval.
- The reliability of AI in providing accurate clinical information is a growing concern.
Purpose of the Study:
- To evaluate the accuracy of ChatGPT in responding to questions based on Society for Neuroscience in Anesthesiology and Critical Care (SNACC) clinical guidelines.
- To assess ChatGPT's ability to incorporate high-quality recommendations (HQRs) from SNACC guidelines on stroke and spine surgery.
- To identify potential inaccuracies or harmful recommendations provided by ChatGPT.
Main Methods:
- Four neuroanesthesiologists independently reviewed ChatGPT responses to 52 HQRs from three SNACC guidelines.
- Assessments included the presence of HQRs, incorrect references, and adherence to guidelines.
- ChatGPT's responses to generic and targeted questions were analyzed.
Main Results:
- Reviewer agreement on HQRs varied widely (0%–100%).
- Only 8% of HQRs were consistently identified after generic questions; this rose to 44% after targeted questions.
- Potentially harmful recommendations were found, and ChatGPT failed to cite the SNACC guidelines.
Conclusions:
- ChatGPT responses require human interpretation and are not consistently accurate regarding HQRs.
- Even with targeted questioning, less than 50% of HQRs were reliably identified.
- Current AI chatbots like ChatGPT are not recommended for clinical decision-making due to reliability concerns.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
08:11Surgical Training for the Implantation of Neocortical Microelectrode Arrays Using a Formaldehyde-fixed Human Cadaver Model
Published on: November 19, 2017