Related Experiment Video
Updated: Jan 12, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance of Large Language Models in Analyzing Common Hypertension Scenarios
Jaleh Zand1, Jing Miao1, Musab S Hommos2
1Division of Nephrology and Hypertension (J.Z., J.M., G.L.S., S.J.T., W.C., V.D.G., Z.M.Z.), Mayo Clinic, Rochester, MN.
Insights
Large language models (LLMs) show potential for aiding hypertension management, with GPT-4 demonstrating the highest accuracy and safety among tested models. However, current LLMs are not yet superior to expert recommendations, necessitating human oversight in clinical applications.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Decision Support Systems
- Cardiovascular Disease Management
Background:
- Hypertension is a leading cause of cardiovascular mortality with suboptimal control rates.
- Large language models (LLMs) offer potential for improving hypertension management by assisting clinical decision-making.
- The reliability of LLMs for guideline-driven medical tasks requires thorough evaluation.
Purpose of the Study:
- To assess the accuracy and safety of hypertension management recommendations generated by three distinct LLMs.
- To compare LLM performance against expert-generated recommendations for hypertension care.
- To determine the guideline concordance and reliability of LLM-based clinical advice.
Main Methods:
- Fifty-one clinical vignettes were developed for hypertension management scenarios.
- Responses were generated by three LLMs (GPT-4, Gemini, MedLM) and a hypertension expert.
- Blinded reviewers evaluated responses for accuracy, safety, and source identification.
Main Results:
- GPT-4 achieved the highest accuracy (83%) and safety (86%) among LLMs, yet remained below expert performance (92% accuracy, 93% safety).
- Gemini (64% accuracy, 73% safety) and MedLM (35% accuracy, 39% safety) showed significantly lower performance.
- GPT-4 exhibited greater guideline concordance (46%) compared to other LLMs but lagged behind expert recommendations (68%).
Conclusions:
- GPT-4 shows promise for supporting hypertension management due to its closer alignment with expert decisions.
- Current LLM capabilities in hypertension management are inferior to those of human experts.
- Continuous human-in-the-loop supervision is crucial for the safe and effective deployment of LLMs in clinical settings.
Background:
Hypertension, the leading cause of cardiovascular mortality, remains suboptimally controlled. Large language models (LLMs) could improve hypertension control by augmenting clinical decision-making, but their reliability for guideline-driven tasks is unverified. This study evaluated the accuracy and safety of hypertension management recommendations generated by 3 LLMs.
Methods:
Fifty-one vignettes were constructed and submitted to the LLMs (GPT-4, Gemini, Medical Large Language Model [by Google; MedLM]) and a hypertension expert to generate the responses. Three blinded reviewers rated each response on a 4-point accuracy scale, a binary safety (safe/unsafe) scale, and attempted to identify the source (LLM versus expert) providing the response.
Results:
GPT-4 had the highest accuracy (83%) and safety (86%) scores among LLMs but remained inferior to expert responses (92% accuracy, 93% safety). Gemini and MedLM performed significantly worse (accuracy: 64% and 35%; safety: 73% and 39%, respectively). GPT-4 generated the most guideline-concordant responses (46%) among the 3 LLMs (Gemini 35%, MedLM 14%) but was lower than expert responses (68%). Interrater reliability for accuracy ratings was higher for LLM-generated responses (GPT-4 [intraclass correlation coefficient, 0.30], Gemini [intraclass correlation coefficient, 0.61], and MedLM [intraclass correlation coefficient, 0.58]), with lower agreement for expert responses (intraclass correlation coefficient, 0.23). A similar pattern was observed for safety and source discrimination ratings. The agreement was strongest for safety assessments and weakest for source discrimination.
Conclusions:
Among the 3 tested LLMs, GPT-4 demonstrated closer agreement to expert decisions, thereby showing greater potential for supporting hypertension management. Despite their potential, current LLM versions are inferior to expert recommendations. Human-in-the-loop supervision remains essential when deploying LLMs for clinical decision-making.
Related Concept Videos
Hypertension III: Clinical Manifestations and Diagnostic Studies
Errors occurring during blood pressure monitoring
Several factors...
Hypertension I: Introduction
Hypertension V: Nursing Management
Pre-Procedural Guidelines for Assessing Blood Pressure
Hypertension and Regulation of Blood Pressure

