Adversarial prompt and fine-tuning attacks threaten medical large language models

Yifan Yang1,2, Qiao Jin1, Furong Huang2

  • 1National Library of Medicine (NLM), National Institutes of Health (NIH), Bethesda, MD, USA.

Nature Communications
|October 9, 2025
PubMed

Insights

Large Language Models (LLMs) in healthcare are vulnerable to adversarial attacks like prompt injections and poisoned data. Detecting shifts in model weights after attacks is key to developing defenses for safe medical AI.

Area of Science:

  • Artificial Intelligence in Medicine
  • Cybersecurity in Healthcare
  • Natural Language Processing

Background:

  • Large Language Models (LLMs) show potential in revolutionizing healthcare, aiding diagnostics and patient care.
  • However, LLMs are susceptible to adversarial attacks, posing risks in sensitive medical applications.

Purpose of the Study:

  • To investigate the vulnerability of LLMs to prompt injection and data poisoning attacks.
  • To assess LLM security across medical tasks including disease prevention, diagnosis, and treatment.

Main Methods:

  • Utilized real-world patient data to test both open-source and proprietary LLMs.
  • Simulated adversarial attacks including prompt injections and fine-tuning with poisoned samples.

Main Results:

  • Demonstrated significant vulnerability of LLMs to malicious manipulation across various medical tasks.
  • Observed subtle shifts in fine-tuned model weights due to poisoned data, even without major performance degradation.

Conclusions:

  • LLMs in healthcare require robust security measures against adversarial attacks.
  • Identifying weight shifts offers a potential method for detecting and mitigating attacks on medical AI systems.