Related Experiment Video
Updated: Jun 20, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Memorization in large language models in medicine prevalence characteristics and implications
Anran Li1, Lingfei Qian1, Mengmeng Du2
1Department of Biomedical Informatics and Data Science, School of Medicine, Yale University, New Haven, CT, USA.
Abstract:
Large Language Models (LLMs) have demonstrated significant potential in medicine, with many studies adapting them through continued pretraining or fine-tuning on medical data. However, a key question remains: to what extent do LLMs memorize medical training data-that is, recall or regenerate content seen during continued pretraining or fine-tuning. In this work, we investigate memorization of LLMs in medicine, assessing its prevalence (frequency), characteristics (what is memorized), volume (how much), and potential downstream impacts. We systematically analyze common adaptation scenarios: (1) continued pretraining on medical corpora, (2) fine-tuning on standard medical benchmarks, and (3) fine-tuning on real-world clinical data, including over 13,000 unique inpatient records from Yale New Haven Health System. The results demonstrate that memorization is prevalent and significantly higher than that in the general domain. Memorization has distinct characteristics during continued pretraining and fine-tuning, and it is persistent: up to 87% of content memorized during continued pretraining remains after fine-tuning. Memorization can be categorized into three types: beneficial (e.g., accurate recall of clinical guidelines), uninformative (e.g., templated language), and harmful (e.g., sensitive clinical content). We offer practical recommendations to facilitate beneficial memorization, minimize uninformative memorization, and mitigate harmful memorization to protect patient privacy and improve medical utility.
Related Concept Videos
Introduction to Language of Pathophysiology l
Introduction to Language of Pathophysiology ll
Language and Cognition
Multicompartment Models: Overview
These models offer a more comprehensive representation of drug behavior in the body than one-compartment models. They accommodate the complexity of drug distribution,...
Guidelines for Nursing Documentation I
Factual:
The following points emphasize the significance of upholding accurate and unbiased documentation in healthcare.
Mnemonic Devices
Acronyms
Acronyms are created by using the initial letters of a series of words to form a new word or phrase. This approach condenses complex information into a single, memorable entity. For example,...