Related Experiment Video
Updated: Jun 20, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Memorization in large language models in medicine prevalence characteristics and implications
Anran Li1, Lingfei Qian1, Mengmeng Du2
1Department of Biomedical Informatics and Data Science, School of Medicine, Yale University, New Haven, CT, USA.
Large Language Models (LLMs) in medicine frequently memorize training data, recalling both beneficial and harmful content. This study analyzes LLM memorization across different medical adaptation scenarios.
Area of Science:
- Artificial Intelligence in Medicine
- Natural Language Processing
- Medical Informatics
Background:
- Large Language Models (LLMs) show promise in medicine, often adapted via continued pretraining or fine-tuning on medical datasets.
- A critical concern is the extent to which these LLMs memorize their medical training data.
Purpose of the Study:
- To investigate the prevalence, characteristics, volume, and impact of LLM memorization in medical training data.
- To analyze memorization across various medical adaptation scenarios.
Main Methods:
- Systematic analysis of LLMs adapted through continued pretraining on medical corpora.
- Evaluation of LLMs fine-tuned on standard medical benchmarks.
- Assessment of LLMs fine-tuned on real-world clinical data (over 13,000 inpatient records).
Main Results:
- Memorization in medical LLMs is prevalent and significantly higher than in general-domain LLMs.
- Memorization exhibits distinct characteristics during pretraining versus fine-tuning.
- Up to 87% of memorized content from pretraining persists after fine-tuning.
Conclusions:
- Memorization is categorized as beneficial (e.g., guidelines), uninformative (e.g., templates), or harmful (e.g., sensitive data).
- Recommendations are provided to leverage beneficial memorization, reduce uninformative memorization, and mitigate harmful memorization for privacy and utility.
Related Concept Videos
Introduction to Language of Pathophysiology l
Introduction to Language of Pathophysiology ll
Language and Cognition
Multicompartment Models: Overview
These models offer a more comprehensive representation of drug behavior in the body than one-compartment models. They accommodate the complexity of drug distribution,...
Guidelines for Nursing Documentation I
Factual:
The following points emphasize the significance of upholding accurate and unbiased documentation in healthcare.
Mnemonic Devices
Acronyms
Acronyms are created by using the initial letters of a series of words to form a new word or phrase. This approach condenses complex information into a single, memorable entity. For example,...