Related Experiment Video
Updated: May 24, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating base and retrieval augmented LLMs with document or online support for evidence based neurology
Lars Masanneck1, Sven G Meuth2, Marc Pawlitzki2
1Department of Neurology, Medical Faculty and University Hospital Düsseldorf, Heinrich Heine University Düsseldorf, Düsseldorf, Germany. lars.masanneck@med.uni-duesseldorf.de.
Large language models (LLMs) and retrieval-augmented generation (RAG) systems show varied accuracy in managing evidence-based neurology information. While RAG improves performance, further development is crucial for safe clinical use.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Medicine
- Clinical Decision Support
Background:
- Managing evidence-based medical information is complex and growing.
- Large language models (LLMs) offer potential solutions for information synthesis.
- Retrieval-augmented generation (RAG) aims to enhance LLM accuracy with external data.
Purpose of the Study:
- To evaluate the performance of LLMs and RAG systems in answering clinical questions based on neurology guidelines.
- To compare document-enabled and online-enabled RAG systems against base LLM performance.
- To identify limitations and areas for improvement in RAG systems for medical applications.
Main Methods:
- Tested LLMs and RAG systems (document- and online-enabled) on 130 questions derived from 13 neurology guidelines.
- Assessed accuracy and identified types of errors, including potentially harmful responses.
- Categorized questions into knowledge-based and case-based for comparative analysis.
Main Results:
- Significant variability in performance was observed across different LLM and RAG configurations.
- RAG systems demonstrated improved accuracy over base LLMs but still generated inaccuracies.
- RAG systems underperformed on case-based clinical scenarios compared to knowledge-based questions.
Conclusions:
- Current RAG-enhanced LLMs require substantial refinement for reliable clinical integration.
- Improved accuracy and safety measures are necessary before widespread adoption in healthcare.
- Further research and regulatory oversight are essential for responsible implementation of AI in clinical practice.
More Related Videos
12:55Multimodal Protocol for Assessing Metacognition and Self-Regulation in Adults with Learning Difficulties
Published on: September 27, 2020
08:05Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
Published on: June 30, 2020