Related Experiment Video
Updated: Jul 12, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
K-anonymity decay in multi-turn clinical large language model conversations
James Weatherhead1, Azra Hasan2, Jake Weatherhead3
1Graduate School of Biomedical Sciences, The University of Texas Medical Branch at Galveston, Galveston, TX, United States.
Per-prompt de-identification in clinical AI conversations may not protect patient privacy. Cumulative disclosures across multiple turns increase re-identification risk, even without direct identifiers.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Health Data Privacy
Background:
- Per-prompt de-identification is standard for clinical AI conversations.
- The Health Insurance Portability and Accountability Act (HIPAA) Safe Harbor method requires removing 18 identifier categories per disclosure.
- Current methods do not assess cumulative re-identification risk across conversational turns.
Purpose of the Study:
- To quantify the re-identification risk in multi-turn clinical AI conversations.
- To evaluate the effectiveness of per-prompt de-identification against cumulative quasi-identifier disclosure.
- To highlight limitations in current privacy protection methods for clinical AI.
Main Methods:
- Simulated progressive quasi-identifier disclosure based on clinical case presentations.
- Modeled disclosure against a synthetic electronic health record cohort.
- Quantified patient re-identification risk by tracking proximity to the small-cell threshold (k < 5).
Main Results:
- 79.9% of simulated patients fell below the k-anonymity threshold by the end of disclosure sequences.
- A median of seven disclosure steps were needed to reach the threshold; this decreased to four steps when rare attributes were disclosed first.
- Cumulative quasi-identifier profiles degraded k-anonymity below safety thresholds, even without direct identifier disclosure.
Conclusions:
- Per-prompt de-identification is insufficient for multi-turn clinical AI and large language model conversations.
- Clinicians lack real-time tools to assess cumulative re-identification risk.
- Existing HIPAA Safe Harbor provisions may not adequately address the evolving privacy challenges in AI-driven healthcare.
Related Concept Videos
Language and Cognition
Nonsense-mediated mRNA Decay
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...
Nonsense-mediated mRNA Decay
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...
Actor-Observer Effect
Parameters Affecting Nonlinear Elimination: Zero-Order Input, First-Order Absorption and Two-Compartment Model
When a drug is administered through a constant intravenous infusion and eliminated via nonlinear pharmacokinetics, it follows zero-order input. For example, oral drugs undergo first-order absorption upon administration and are eliminated through nonlinear pharmacokinetics.
In the case of subcutaneously administered drugs,...
Woodward–Hoffmann Selection Rules and Microscopic Reversibility