Related Experiment Video
Updated: Aug 6, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Benchmarking large language models for de-identification of electronic health record notes
Omkar Panchal1, Nai-Wen Chang2, Zi-Rui Zhao3
1CGD Health Pvt Ltd, Mumbai, Maharashtra, India.
BMJ Health & Care Informatics
|July 21, 2026
Summary
Large language models (LLMs) show promise for de-identifying sensitive health information (SHI). Fine-tuned LLMs achieved a high F1 score of 0.9447, outperforming traditional methods, but variability across datasets presents challenges.
Area of Science:
- Natural Language Processing
- Clinical Informatics
- Health Data Security
Background:
- The increasing use of large language models (LLMs) in processing clinical text necessitates robust de-identification methods.
- Evaluating LLM capabilities for identifying sensitive health information (SHI) is crucial for ensuring data privacy.
- Existing de-identification techniques require benchmarking against advanced LLM-based approaches.
Purpose of the Study:
- To conduct a comprehensive benchmarking analysis of LLM-based, rule-based, and hybrid de-identification methods.
- To evaluate the performance and robustness of various de-identification models across diverse datasets.
- To compare the effectiveness of different LLM fine-tuning strategies for SHI detection.
Main Methods:
- Utilized five diverse datasets (i2b2-2006, MIMIC-2008, i2b2-2014, i2b2-2016, OpenDeID v1) from multiple countries.
- Developed three baseline and eight LLM-based de-identification models.
- Implemented nine experimental settings to assess cross-dataset performance and model robustness.
Main Results:
- The best baseline model, trained on a combined corpus, achieved a strict F1 micro-average score of 0.8172.
- Supervised fine-tuning of LLMs using a combined corpus configuration yielded the highest performance with a strict F1 score of 0.9447.
- Performance varied significantly across heterogeneous electronic health record sources.
Conclusions:
- Harmonizing corpora is essential for standardized data formatting and SHI management, improving de-identification reliability.
- Fine-tuned LLMs demonstrate superior accuracy in de-identification but exhibit performance variability.
- Addressing inconsistencies in electronic health record data is critical for the ethical and technical deployment of LLMs in healthcare.