Automatic de-identification of electronic medical records using token-level and character-level conditional random

Zengjian Liu1, Yangxin Chen2, Buzhou Tang1

  • 1Key Laboratory of Network Oriented Intelligent Computation, Harbin Institute of Technology Shenzhen Graduate School, Shenzhen 518055, China.

Summary

This study presents a hybrid system for de-identifying clinical data, achieving top ranks in the 2014 i2b2 NLP challenge. The machine learning and rule-based approach effectively removes protected health information (PHI) from electronic medical records (EMRs).