Related Experiment Video
Updated: Jul 3, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Almanac - Retrieval-Augmented Language Models for Clinical Medicine
Cyril Zakka1, Rohan Shad2, Akash Chaurasia3
1Department of Cardiothoracic Surgery, Stanford Medicine, Stanford, CA.
Large language models (LLMs) show promise in medicine but can err. Almanac, an LLM with medical data access, demonstrated improved accuracy and safety over standard LLMs in clinical question answering.
Area of Science:
- Artificial Intelligence
- Clinical Medicine
- Natural Language Processing
Background:
- Large language models (LLMs) exhibit zero-shot capabilities for various natural language tasks.
- Despite potential in clinical medicine, LLM adoption is hindered by factual inaccuracies and safety concerns.
Purpose of the Study:
- To evaluate Almanac, an LLM framework with retrieval from curated medical resources, for medical guideline and treatment recommendations.
- To compare Almanac's performance against standard LLMs (ChatGPT-4, Bing, Bard) using clinical questions.
Main Methods:
- A panel of eight board-certified clinicians and two health care practitioners evaluated LLM responses.
- The evaluation used a novel dataset of 314 clinical questions across nine medical specialties.
- Responses from Almanac, ChatGPT-4, Bing, and Bard were compared.
Main Results:
- Almanac demonstrated significant improvements in factuality and completeness compared to standard LLMs.
- Almanac also showed enhanced user preference and adversarial safety.
- Performance was assessed across nine medical specialties.
Conclusions:
- LLMs accessing domain-specific corpora hold potential for clinical decision-making.
- Rigorous testing of LLMs is crucial before clinical deployment to address limitations.
- Findings supported by the National Institutes of Health, National Heart, Lung, and Blood Institute.
More Related Videos
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
10:15Utilizing Repetitive Transcranial Magnetic Stimulation to Improve Language Function in Stroke Patients with Chronic Non-fluent Aphasia
Published on: July 2, 2013
Related Concept Videos
Retrieval
Recall involves accessing information without cues, such as during an essay test, where individuals must retrieve facts and concepts from memory unaided. Another example is remembering the name of a colleague...
Clinical Trials: Overview
Clinical Trials
There are four phases in a clinical trial. A phase one...