Related Experiment Video
Updated: Feb 28, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
When Agentic LLMs Trust Poisoned Tools: Vulnerability of Clinical LLMs to Adversarial Guidelines
Mahmud Omar1, Alon Gorenshtien1, Yiftach Barash2
1Icahn School of Medicine at Mount Sinai.
Agentic large language models (LLMs) struggle to reject modified medical guidelines, frequently selecting incorrect information. This vulnerability poses risks, especially when AI agents act as primary health gatekeepers.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Informatics
- Natural Language Processing
Background:
- Agentic large language models (LLMs) are increasingly integrated with external tools and data sources.
- The reliability of LLMs in selecting accurate information from potentially compromised sources remains an open question.
Purpose of the Study:
- To evaluate the ability of 21 LLMs to discern authentic medical guidelines from adversarially modified versions.
- To identify failure rates and influencing factors in LLM tool selection for medical decision-making.
Main Methods:
- 21 LLMs were tested on 500 physician-validated medical vignettes across 12 domains.
- Models chose between authentic and sham (adversarially modified) guideline excerpts.
- A total of 10,500 agentic decisions were analyzed, considering presentation order.
Main Results:
- LLMs selected the sham, incorrect guideline in 40.6% of cases, achieving 59.4% accuracy.
- Failure rates were highest for safety-critical modifications (54.2%-61.7%), including altered warnings, allergy information, contraindications, and dosing.
- Model choices were heavily influenced by presentation bias, favoring the first presented option.
Conclusions:
- Agentic LLM guideline selection is vulnerable to poisoned sources, necessitating safeguards before clinical deployment.
- Independent verification and ranking mechanisms are crucial for ensuring AI agent reliability in healthcare.
- Low-resource settings relying on AI agents face heightened risks from unreliable tools.
Related Concept Videos
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Improving Translational Accuracy
Improving Translational Accuracy
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Pharmaceutical Poisoning: Potential Scenarios
Language and Cognition
