Related Experiment Video
Updated: Jun 14, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
519
DIRI: Adversarial Patient Reidentification with Large Language Models for Evaluating Clinical Text Anonymization
John X Morris1, Thomas R Campion1, Sri Laasya Nutheti1
1Cornell Tech, New York, NY.
Summary
Current deidentification methods fail to fully protect patient privacy in clinical notes. An adversarial large language model (LLM) approach successfully re-identified 9% of notes, revealing weaknesses in existing tools.
Area of Science:
- Biomedical Informatics
- Natural Language Processing
- Data Privacy
Background:
- Sharing protected health information (PHI) is vital for biomedical research.
- Deidentification is crucial for removing PHI from clinical text before data distribution.
- Current deidentification methods are often evaluated on limited datasets, potentially overestimating real-world performance.
Purpose of the Study:
- To develop and evaluate a novel adversarial method using a large language model (LLM) to re-identify patients from de-identified clinical notes.
- To assess the effectiveness of state-of-the-art deidentification tools against a re-identification attack.
- To highlight limitations in current deidentification technologies and provide a tool for iterative improvement.
Main Methods:
- Developed an adversarial approach using a large language model (LLM) for re-identification.
- Introduced a De-Identification/Re-Identification (DIRI) method to evaluate deidentification tool performance.
- Tested the method on clinical data from Weill Cornell Medicine anonymized using Philter, BiLSTM-CRF, and ClinicalBERT.
Main Results:
- The LLM-based re-identification tool successfully re-identified 9% of clinical notes, even those processed by the most effective deidentification tool (ClinicalBERT).
- This demonstrates significant weaknesses in current deidentification technologies.
- The DIRI method provides a robust evaluation framework for deidentification tools.
Conclusions:
- Existing deidentification technologies exhibit significant vulnerabilities.
- The developed LLM-based re-identification method can effectively challenge and expose these weaknesses.
- Continuous improvement and novel approaches are necessary to ensure robust patient privacy in biomedical research.

