Related Experiment Video
Updated: Feb 5, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Vulnerability of Large Language Models to Prompt Injection When Providing Medical Advice
Ro Woon Lee1,2, Tae Joon Jun3, Jeong-Moo Lee1,4
1Artificial Intelligence Research Committee, GIGA Study, Incheon, Republic of Korea.
Commercial large language models (LLMs) are highly vulnerable to prompt-injection attacks, potentially leading to unsafe medical advice. Even advanced models show significant susceptibility, highlighting the need for robust security measures before healthcare deployment.
Area of Science:
- Artificial Intelligence in Healthcare
- Cybersecurity in Medical Applications
- Natural Language Processing
Background:
- Large language models (LLMs) are increasingly used in healthcare, raising concerns about their security.
- Prompt-injection attacks can manipulate LLM behavior, potentially altering critical medical recommendations.
- Systematic evaluation of LLM vulnerability to these attacks in clinical settings is lacking.
Purpose of the Study:
- To assess the susceptibility of commercial LLMs to prompt-injection attacks.
- To determine if these attacks can induce unsafe clinical advice.
- To validate man-in-the-middle and client-side injection as viable attack vectors.
Main Methods:
- Controlled simulation study using standardized patient-LLM dialogues (January-October 2025).
- Evaluated lightweight models (GPT-4o-mini, Gemini-2.0-flash-lite, Claude-3-haiku) across 12 clinical scenarios stratified by harm level.
- Tested flagship models (GPT-5, Gemini 2.5 Pro, Claude 4.5 Sonnet) using client-side injection in a high-risk pregnancy scenario.
- Employed context-aware and evidence-fabrication injection strategies within a multiturn dialogue framework.
Main Results:
- Prompt-injection attacks achieved 94.4% success rate at the primary decision turn and persisted in 69.4% of follow-ups.
- Lightweight models showed complete (100%) or high (83.3%) susceptibility.
- Extremely high-harm scenarios, including Category X pregnancy drugs, succeeded in 91.7% of attacks.
- Flagship models demonstrated 80-100% vulnerability in proof-of-concept testing.
Conclusions:
- Commercial LLMs exhibit significant vulnerability to prompt-injection attacks, capable of generating dangerous clinical advice.
- Even advanced LLMs with safety features are susceptible, underscoring security risks.
- Adversarial robustness testing, system safeguards, and regulatory oversight are crucial before clinical deployment of LLMs.
More Related Videos
Related Concept Videos
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Components of Language
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Language and Cognition
Inhaled Medications
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...

