Related Experiment Video
Updated: Jul 9, 2026

Inverse Probability of Treatment Weighting (Propensity Score) using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
A de-identifier for medical discharge summaries
Ozlem Uzuner1, Tawanda C Sibanda, Yuan Luo
1University at Albany, State University of New York, Draper 114, Albany, NY 12222, USA. ouzuner@albany.edu
We developed Stat De-id, a machine learning tool that effectively removes personal health information (PHI) from clinical discharge summaries, even with complex language. This de-identification method achieves high accuracy, improving data usability for research.
Area of Science:
- Medical Informatics
- Natural Language Processing
- Data Privacy
Background:
- Clinical records contain valuable research data but are protected by privacy regulations due to personal health information (PHI).
- De-identifying clinical records is crucial for data sharing but challenging due to linguistic complexities like fragmented language, domain-specific terms, and ambiguous PHI.
- Existing tools often struggle with the unique characteristics of medical texts, such as discharge summaries.
Purpose of the Study:
- To develop and evaluate an automated de-identification system for clinical discharge summaries.
- To assess the effectiveness of a novel approach emphasizing local context for de-identification.
- To compare the performance of the proposed system against existing rule-based and named entity recognition (NER) methods.
Main Methods:
- Developed 'Stat De-id,' a de-identifier utilizing support vector machines and a detailed representation of local context.
- Evaluated Stat De-id on clinical discharge summaries, considering challenges like out-of-vocabulary words and ambiguous PHI.
- Compared Stat De-id's performance (F-measure) against a rule-based approach, SNoW, IdentiFinder, and a Conditional Random Field De-identifier (CRFD).
Main Results:
- Stat De-id achieved a high F-measure of 97% for de-identifying personal health information (PHI).
- The system demonstrated superior performance over a rule-based approach (85% F-measure) and NER systems like IdentiFinder (36% F-measure).
- Stat De-id (88% F-measure) outperformed CRFD, indicating that enhancing local context representation is more beneficial than solely adding global context for de-identification in fragmented medical text.
Conclusions:
- A de-identifier leveraging support vector machines and robust local context representation can effectively de-identify clinical discharge summaries.
- This approach is particularly effective in handling linguistic challenges present in medical texts, such as fragmented language and ambiguous PHI.
- Prioritizing the representation of local context in de-identification systems offers significant advantages over methods relying heavily on global context or simpler contextual representations.
Related Concept Videos
Discharge Summary Forms
Here's a detailed look at the key components and guidelines for preparing a discharge summary:
Flow Sheet
Here's a closer look at the examples of flowsheets commonly used by nurses:
Graphic Sheet Documentation:
Methods of Documentation I: Source-Oriented Records
In an SOR, each discipline involved in patient care maintains a separate medical record section. This record-keeping method enables easy tracking of patient progress and ensures healthcare staff have access to up-to-date information.
Key Attributes include the following:
Data Reporting and Recording
Methods of Documentation VII: EMR
Methods of Documentation V: CBE
In CBE, healthcare professionals establish predefined standards of practice that define what constitutes...
