Related Experiment Video
Updated: Mar 5, 2026

07:50
A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
16.6K
Semi-automated De-identification of German Content Sensitive Reports for Big Data Analytics
Hannes Seuss1, Peter Dankerl1, Matthias Ihle2
1Department of Radiology, University Hospital Erlangen, Friedrich Alexander Universität (FAU) Erlangen-Nürnberg, Erlangen, Germany.
Summary
Developing a semi-automated tool for de-identifying medical reports significantly improves data security for collaborations. Training the software with approximately 100 reports ensures reliable detection of sensitive information.
Area of Science:
- Medical Informatics
- Data Security
- Natural Language Processing
Background:
- Collaborative research across institutions necessitates robust data security measures.
- Selective de-identification of sensitive information in medical reports is crucial for data sharing.
- The increasing volume of 'Big Data' in healthcare amplifies the need for automated de-identification solutions.
Purpose of the Study:
- To develop and evaluate a semi-automated tool for de-identifying sensitive content in various medical reports.
- To assess the tool's performance in detecting direct and indirect identifiers, medical terms, and filler words.
- To determine the optimal amount of training data required for reliable de-identification.
Main Methods:
- A semi-automated de-identification tool was developed and tested on diverse medical reports (pathology, medical, operation, radiology).
- The tool was evaluated natively and after iterative retraining with increasing sets of manually edited reports (25, 50, 100, 250, 500, 1000).
- Performance was measured by sensitivity and specificity in detecting different categories of sensitive data.
Main Results:
- Native detection rates were 61.3% for direct and 80.8% for indirect identifiers.
- After training with approximately 100 reports, detection rates for direct and indirect identifiers reached 99.5% and 97.2%, respectively.
- Falsely flagged medical terms decreased from 5.3% (native) to approximately 3-4% after training, with minimal flagging of filler words.
Conclusions:
- Continuous training significantly enhances the performance of the developed de-identification tool.
- Training with around 100 edited reports enables reliable detection and labeling of sensitive data across different medical report types.
- While the tool is highly effective, a final review by authorized personnel remains essential for complete data security.

