Related Experiment Video
Updated: May 2, 2026

Quantifying Intermembrane Distances with Serial Image Dilations
Published on: September 28, 2018
Measuring the gap: correlating synthetic-to-real drift with PHI de-identification performance
Joseph Cornelius1,2, Fabio Rinaldi3
1Dalle Molle Institute for Artificial Intelligence Research (IDSIA USI-SUPSI), Via la Santa 1, Lugano-Viganello, CH-6962, Ticino, Switzerland. joseph.cornelius@idsia.ch.
Synthetic clinical notes generated by large language models (LLMs) aid de-identification in low-resource settings, but their utility depends on data source and quality control. Drift estimation can improve synthetic data alignment.
Area of Science:
- Natural Language Processing
- Clinical Informatics
- Machine Learning
Background:
- Electronic health records (EHRs) require de-identification for patient privacy.
- Publicly available de-identification training data are scarce and often lack consistent documentation styles.
- Large language models (LLMs) offer a potential solution for generating synthetic clinical notes.
Purpose of the Study:
- To evaluate the impact of lexical and semantic drift on protected health information (PHI) tagger performance.
- To assess the utility of LLM-generated synthetic clinical notes for training de-identification models.
- To determine if synthetic data reflects real-world clinical note distributions.
Main Methods:
- Generated synthetic clinical notes using five generator LLMs and one judge LLM.
- Fine-tuned de-identification models on real, synthetic, and mixed corpora.
- Evaluated model performance on three external benchmarks using a harmonized label schema.
- Assessed correlation between drift measures and out-of-distribution F1 score.
Main Results:
- Models trained on broad, clinically relevant data sources outperformed those trained on legal or narrowly synthetic data.
- Synthetic data, despite lacking some real-world distributional properties, proved useful in low-resource scenarios.
- Compact distributional and embedding-based drift measures showed moderate correlation with out-of-distribution F1 score.
Conclusions:
- LLM-generated synthetic clinical notes can augment scarce real-world data for de-identification tasks.
- Careful selection of training data sources and quality control through drift estimation are crucial for maximizing synthetic data utility.
- Drift estimation offers a practical method for improving synthetic data quality and alignment in clinical text de-identification.
Related Concept Videos
Calibration Curves: Correlation Coefficient
Deindividuation
¹³C NMR: Distortionless Enhancement by Polarization Transfer (DEPT)
Difference from Background: Limit of Detection
The LOD indicates the presence or absence...
Self-Discrepancy and Its Effects
Detection of Gross Error: The Q Test

