Related Experiment Video
Updated: Jul 3, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
1.3K
Cardiology knowledge assessment of retrieval-augmented open versus proprietary large language models
Constantine Tarabanis1, Shaan Khurshid2,3,4, Areti Karamanou5
1Cardiology Division, Heart and Vascular Institute, Mass General Brigham, Boston, Massachusetts, United States of America.
PLOS Digital Health
|March 12, 2026
Summary
Open-weight large language models (LLMs) show strong performance in cardiology questions, with some models outperforming human averages. Retrieval-Augmented Generation (RAG) further enhances performance, especially for smaller models.
Area of Science:
- Artificial Intelligence in Medicine
- Cardiovascular Disease Research
- Medical Education Technology
Background:
- Large language models (LLMs) are increasingly used in healthcare, but their efficacy in specialized fields like cardiology requires thorough evaluation.
- Assessing LLM performance on board-style questions is crucial for understanding their potential in clinical decision support and medical education.
Purpose of the Study:
- To evaluate and compare the performance of open-weight and proprietary LLMs on cardiology board-style questions.
- To benchmark LLM performance against human performance and assess the impact of Retrieval-Augmented Generation (RAG).
Main Methods:
- 14 LLMs (6 open-weight, 8 proprietary) were tested on 449 multiple-choice questions from the American College of Cardiology Self-Assessment Program (ACCSAP).
- Retrieval-Augmented Generation (RAG) was implemented using a knowledge base of 123 guideline and textbook documents.
- Accuracy was measured as the percentage of correctly answered questions.
Main Results:
- The open-weight model DeepSeek R1 achieved the highest accuracy (86.9%), surpassing proprietary models and the human average (78%).
- GPT 4o (80.9%) and OpenEvidence (81.3%) showed comparable performance to each other.
- All models improved with RAG, with open-weight models like Mistral Large 2 performing similarly to proprietary models like GPT 4o.
Conclusions:
- Open-weight LLMs can achieve performance comparable to or exceeding proprietary models in cardiovascular medicine, with or without RAG.
- RAG significantly benefits all models, particularly smaller open-weight ones, enhancing their utility in clinical applications.
- Open-weight LLMs offer a viable, transparent, and configurable alternative for clinical use, potentially at a lower cost.

