Related Experiment Video
Updated: Apr 8, 2026

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Automatic de-identification of electronic medical records using token-level and character-level conditional random
Zengjian Liu1, Yangxin Chen2, Buzhou Tang1
1Key Laboratory of Network Oriented Intelligent Computation, Harbin Institute of Technology Shenzhen Graduate School, Shenzhen 518055, China.
This study presents a hybrid system for de-identifying clinical data, achieving top ranks in the 2014 i2b2 NLP challenge. The machine learning and rule-based approach effectively removes protected health information (PHI) from electronic medical records (EMRs).
Area of Science:
- Natural Language Processing (NLP)
- Clinical Informatics
- Data Privacy
Background:
- De-identification of protected health information (PHI) is crucial for public data release.
- Electronic medical records (EMRs) contain sensitive patient data requiring robust de-identification methods.
- The 2014 i2b2 NLP challenge focused on de-identification of clinical text.
Purpose of the Study:
- To develop and evaluate a hybrid system for clinical data de-identification.
- To improve the accuracy of identifying and removing PHI from EMRs.
- To compete in the de-identification track of the 2014 i2b2 challenge.
Main Methods:
- A hybrid system combining machine learning and rule-based approaches was developed.
- Two token-level and character-level conditional random fields (CRFs) were used for PHI identification.
- A rule-based classifier and merging rules were integrated for enhanced accuracy.
Main Results:
- The system achieved top-ranked micro F-scores on the i2b2 corpus: 94.64% (token), 91.24% (strict), and 91.63% (relaxed).
- Further improvements with refined dictionaries yielded F-scores of 94.83% (token), 91.57% (strict), and 91.95% (relaxed).
- The hybrid system demonstrated high performance in the 2014 i2b2 de-identification challenge.
Conclusions:
- The proposed hybrid system is effective for de-identifying clinical data.
- The combination of machine learning and rule-based methods enhances PHI removal accuracy.
- The system's performance meets high standards for clinical data privacy and usability.
More Related Videos
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
12:18A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Related Concept Videos
Methods of Documentation VII: EMR
Guidelines and Strategies for Safe Computer Charting
Maintain Confidentiality and Security:
Randomized Experiments
Simple randomization
Simple...