Related Experiment Video
Updated: Aug 6, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Benchmarking large language models for de-identification of electronic health record notes
Omkar Panchal1, Nai-Wen Chang2, Zi-Rui Zhao3
1CGD Health Pvt Ltd, Mumbai, Maharashtra, India.
BMJ Health & Care Informatics
|July 21, 2026
Summary
Large language models (LLMs) show promise for de-identifying sensitive health information (SHI). Fine-tuned LLMs achieved a high F1 score of 0.9447, outperforming traditional methods, but implementation requires addressing data inconsistencies.
Area of Science:
- Natural Language Processing
- Clinical Informatics
- Health Data Privacy
Background:
- The increasing use of large language models (LLMs) in processing clinical text necessitates robust de-identification methods.
- Evaluating LLM capabilities for identifying sensitive health information (SHI) is crucial for reliable clinical text processing.
Purpose of the Study:
- To benchmark LLM-based, rule-based, and hybrid de-identification methods.
- To assess the performance and robustness of various de-identification approaches across diverse datasets.
Main Methods:
- Utilized five international datasets (i2b2-2006, MIMIC-2008, i2b2-2014, i2b2-2016, OpenDeID v1).
- Developed three baseline and eight LLM-based models.
- Implemented nine experimental settings to evaluate cross-dataset performance and model robustness.
Main Results:
- The best baseline model achieved a strict F1 micro-average score of 0.8172 when trained on a combined corpus.
- Supervised fine-tuning of LLMs (setting 9) yielded the highest performance with a strict F1 score of 0.9447.
- Harmonizing corpora improved data standardization and SHI management.
Conclusions:
- Fine-tuned LLMs demonstrate superior accuracy in de-identification.
- Performance variability across diverse electronic health record sources presents technical challenges for LLM implementation.
- Addressing data inconsistencies is vital for the ethical and technical deployment of LLMs in handling sensitive health data.