Related Experiment Video
Updated: May 27, 2025

07:50
A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
15.7K
De-identification of clinical notes with pseudo-labeling using regular expression rules and pre-trained BERT.
Jiyong An1, Jiyun Kim1, Leonard Sunwoo2
1Graduate School of Data Science, Seoul National University, Seoul, South Korea.
BMC Medical Informatics and Decision Making
|February 18, 2025
Summary
This study developed a semi-supervised method to de-identify Korean clinical notes, combining rule-based and machine learning approaches. The combined strategy significantly improved the accuracy of removing personal information from medical text.
Area of Science:
- Medical Informatics
- Natural Language Processing
- Data Privacy
Background:
- De-identification of clinical notes is crucial for medical research using unstructured text data.
- Limited research exists on de-identifying Korean clinical notes.
- Protecting patient privacy in electronic health records is a growing concern.
Purpose of the Study:
- To develop and evaluate an effective de-identification method for Korean clinical notes.
- To improve the utilization of unstructured clinical text data in Korean medical research.
- To address the gap in de-identification techniques for Korean language medical records.
Main Methods:
- Utilized a large dataset from Seoul National University Bundang Hospital, including radiology and non-radiology notes.
- Employed a two-stage de-identification strategy: rule-based (regular expressions) and semi-supervised machine learning (Korean BERT).
- A rule-based approach was refined using expert-annotated data, then used for pseudo-labeling to fine-tune a pre-trained Korean BERT model.
Main Results:
- The rule-based approach achieved 97.2% precision, 93.7% recall, and 96.2% F1 score on radiology notes.
- The semi-supervised KoBERT-NER model, fine-tuned with pseudo-labeled data, reached 96.5% precision, 97.6% recall, and 97.1% F1 score.
- Validation demonstrated high performance in token-level de-identification for both radiology and non-radiology notes.
Conclusions:
- Combining rule-based and semi-supervised machine learning enhances clinical note de-identification performance.
- The proposed method offers an effective solution for de-identifying Korean clinical text.
- Improved de-identification facilitates secure data sharing and analysis in Korean medical research.
Related Concept Videos
Labeling Emotion
91
Emotional labeling is a cognitive process that involves identifying and naming one's emotions, such as anger, fear, happiness, or sadness. It allows individuals to recognize and express their internal emotional states, a critical aspect of emotional regulation and communication. Labeling emotions requires more than mere recognition; it also involves drawing upon memory and contextual cues to understand the current situation and apply a corresponding emotional label. For instance, feeling...
91
Blinding
2.4K
Blinding is a commonly used method of not telling participants which treatment a subject is receiving. Blinding is a critical part of a randomized control trial or RCT. It reduces the bias that affects the results. In an RCT, blinding is used in the form of a placebo. A placebo effect occurs when untreated subjects falsely believe they have received the treatment and report improved symptoms. A placebo or a dummy treatment is administered to subjects to negate the bias caused by such an effect.
2.4K
Deindividuation
26.2K
Deindividuation is a form of social influence on an individual’s behavior such that the individual engages in unusual or non-normal behavior while in a group setting. Why? Because in these group settings, the individual no longer sees themselves as an individual anymore, disinhibiting their behavior and personal restraint.
26.2K

