Related Experiment Video
Updated: Jan 16, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Ensuring Safety and Trust: Analyzing the Risks of Large Language Models in Medicine.
Yifan Yang1,2, Qiao Jin1, Robert Leaman1
1Division of Intramural Research, National Library of Medicine (NLM), National Institutes of Health (NIH), Bethesda, MD 20894, USA.
Large Language Models (LLMs) show potential in healthcare but have significant safety gaps. Current models perform poorly on medical benchmarks, necessitating human oversight and AI safety guardrails for trustworthy medical AI.
Area of Science:
- Artificial Intelligence
- Medical Informatics
- Healthcare Technology
Background:
- Large Language Models (LLMs) demonstrate advanced capabilities, driving interest in their application within healthcare.
- However, a systematic characterization of the risks associated with LLMs in medical contexts is lacking.
- Ensuring safe and trustworthy AI is paramount for clinical adoption.
Purpose of the Study:
- To systematically characterize the risks of LLMs in medical applications.
- To propose a framework for safe and trustworthy medical AI, including five key principles: Truthfulness, Resilience, Fairness, Robustness, and Privacy.
- To introduce the MedGuard benchmark for evaluating LLM performance in healthcare.
Main Methods:
- Development of a comprehensive framework for medical AI safety.
- Creation of the MedGuard benchmark, comprising 1,000 expert-verified medical questions.
- Evaluation of 11 commonly used LLMs against the MedGuard benchmark.
Main Results:
- Current LLMs generally exhibit poor performance across most MedGuard benchmarks, even with safety alignment.
- Performance of LLMs significantly lags behind that of human physicians in medical tasks.
- A notable safety gap exists between current LLMs and the requirements for safe clinical deployment.
Conclusions:
- Despite LLMs' potential, a significant safety gap hinders their widespread adoption in healthcare.
- Human oversight and robust AI safety guardrails are crucial for mitigating risks.
- Further research and development are needed to enhance the safety and reliability of LLMs in medical applications.
More Related Videos
Related Concept Videos
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Language and Cognition
Improving Translational Accuracy
Improving Translational Accuracy
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
Guidelines for Nursing Documentation I
Factual:
The following points emphasize the significance of upholding accurate and unbiased documentation in healthcare.

