Related Experiment Video
Updated: Jul 18, 2025

05:56
Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
2.5K
Web-Based Application Based on Human-in-the-Loop Deep Learning for Deidentifying Free-Text Data in Electronic Medical
Leibo Liu1, Oscar Perez-Concha1, Anthony Nguyen2
1Centre for Big Data Research in Health, University of New South Wales, Sydney, Australia.
Interactive Journal of Medical Research
|August 25, 2023
Summary
This study introduces DEFT, a web-based system for deidentifying electronic medical record text. DEFT enhances privacy by accurately removing personally identifiable information (PII) and improving research data usability.
Area of Science:
- Health Informatics
- Medical Data Privacy
- Natural Language Processing
Background:
- Electronic medical records (EMRs) contain valuable clinical data for research.
- Protecting patient privacy necessitates deidentifying personally identifiable information (PII) in EMR free text.
- Manual deidentification is inefficient, driving the need for automated solutions.
Purpose of the Study:
- Develop an accurate, user-friendly, web-based system for deidentifying EMR free text, named DEFT.
- Enhance adoption in real-world settings through features like interactive learning and collaboration.
- Facilitate secondary use of clinical data while maintaining patient privacy.
Main Methods:
- DEFT utilizes a Bidirectional Long Short-Term Memory-Conditional Random Field (BiLSTM-CRF) deep learning model with RoBERTa embeddings.
- An interactive learning loop with preannotation accelerates manual deidentification.
- The system supports project management, customizable PII types, and user access control.
Main Results:
- DEFT achieved superior performance on the 2014 i2b2 dataset, with microaverage F1-scores of 0.9627.
- Real-world application on clinical notes yielded a microaverage F1-score of 0.9507.
- Preannotation increased manual annotation efficiency by 43%.
Conclusions:
- DEFT provides an accessible solution for deidentifying EMR free text for researchers and data custodians.
- The system's interactive learning loop and user-friendly interface lower the technical barrier for deidentification tasks.
- DEFT facilitates secure and efficient utilization of clinical narrative data.

