Related Experiment Video
Updated: May 15, 2026

07:50
A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Generalizability and comparison of automatic clinical text de-identification methods and resources
Óscar Ferrández1, Brett R South, Shuying Shen
1Department of Biomedical Informatics, University of Utah, Salt Lake City, UT, USA.
AMIA ... Annual Symposium Proceedings. AMIA Symposium
|January 11, 2013
Summary
The Veteran's Health Administration's (VHA) BoB system effectively de-identifies clinical text, balancing patient privacy and data utility. Other systems like MIST and HIDE show high precision but sometimes lower recall for sensitive data.
Area of Science:
- Health Informatics
- Natural Language Processing
- Data Privacy
Background:
- Automated de-identification of clinical text is crucial for protecting patient privacy in electronic health records.
- Evaluating system performance across diverse datasets is essential for ensuring generalizability and portability.
Purpose of the Study:
- To evaluate the performance of the VHA's hybrid best-of-breed (BoB) clinical text de-identification system.
- To compare BoB against two machine learning-based systems, MIST and HIDE.
- To assess the generalizability and portability of de-identification models across different clinical corpora.
Main Methods:
- Evaluation of three automated de-identification systems: BoB, MIST, and HIDE.
- Utilized two distinct clinical corpora: a manually annotated VHA corpus and the 2006 i2b2 de-identification challenge corpus.
- Focused on measuring recall and precision to assess patient privacy protection and document interpretability.
Main Results:
- BoB achieved a recall of 92.6% and precision of 83.6%, demonstrating strong patient privacy protection and competitive interpretability.
- MIST and HIDE exhibited high precision (92.6% and 93.6% respectively) but showed lower recall for sensitive Protected Health Information (PHI) categories.
- Performance varied across document sources, highlighting the importance of model generalizability.
Conclusions:
- The VHA's BoB system offers a robust solution for clinical text de-identification, effectively balancing privacy and utility.
- Machine learning systems like MIST and HIDE provide high precision but require further optimization for comprehensive recall of sensitive data.
- Future research should focus on enhancing model portability and generalizability across diverse healthcare data sources.
