Related Experiment Video
Updated: May 24, 2026

A Bilingual Computational Workflow for Identifying Potential PLK1 Inhibitors in American Sign Language and English
Published on: April 3, 2026
CONORM-DEID: Robustness Evaluation of a Multilingual De-Identification System for Clinical Texts.
Anthony Yazdani1, Alban Bornet1, Hossein Rouhizadeh1
1Department of Radiology and Medical Informatics, University of Geneva, Switzerland.
Clinical text de-identification is crucial for medical research. CONORM-DEID, a multilingual model, shows that synthetic pre-training improves robustness only with significant human data, otherwise it can reduce performance.
Area of Science:
- Natural Language Processing
- Medical Informatics
- Computational Linguistics
Background:
- Clinical narratives hold vital research data but require de-identification to protect patient privacy.
- High recall is essential in de-identification to prevent disclosure of personally identifiable information.
- Multilingual capabilities are needed for diverse clinical research settings.
Purpose of the Study:
- To introduce CONORM-DEID, a multilingual model for clinical text de-identification.
- To evaluate the impact of synthetic pre-training and human-annotated data on model robustness in English and French.
- To compare performance of models trained from scratch versus those with synthetic pre-training.
Main Methods:
- Developed CONORM-DEID, a multilingual de-identification model.
- Conducted experiments comparing models trained from scratch with those using synthetic pre-training.
- Evaluated model robustness across high- and low-data scenarios in English and French.
- Measured performance using F1-scores and precision at 99% recall.
Main Results:
- Models trained from scratch achieved high F1-scores (over 97%).
- Robustness (precision at 99% recall) improved with synthetic pre-training only after substantial human fine-tuning.
- In low-data settings, synthetic pre-training negatively impacted performance, indicating negative transfer.
Conclusions:
- Substantial human fine-tuning is necessary to leverage synthetic pre-training for robust clinical de-identification.
- The effectiveness of synthetic pre-training is context-dependent on the availability of annotated data.
- CONORM-DEID offers a promising approach to multilingual clinical text de-identification, with performance contingent on data augmentation strategies.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Methods of Classification and Identification
Classification of Systems-II
Detection of Gross Error: The Q Test