Related Experiment Video
Updated: Jun 28, 2025

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Impact of Translation on Biomedical Information Extraction: Experiment on Real-Life Clinical Notes
Christel Gérardin1, Yuhan Xiong1,2, Perceval Wajsbürt3
1Institut Pierre Louis d'Epidémiologie et de Santé Publique, Sorbonne Université, Institut National de la Santé et de la Recherche Médicale, Paris, France.
French natural language processing models outperform English models for French medical concept extraction. Native French models are more effective, even with limited annotated data, compared to translated English approaches.
Area of Science:
- Natural Language Processing
- Medical Informatics
- Computational Linguistics
Background:
- Biomedical natural language processing (NLP) typically relies on English models, despite challenges in creating annotated datasets.
- Advancements in machine translation offer potential for using English NLP tools with other languages.
Purpose of the Study:
- To evaluate if English NLP tools, using translated French medical concepts, match the performance of French NLP models trained on French clinical notes.
- To compare the efficacy of native French NLP models versus English NLP models with translation for French medical concept extraction and normalization.
Main Methods:
- A comparative study was conducted using native French-language models and English-language models with translation.
- The native French method involved separate named entity recognition (NER) and normalization steps.
- The English-language method involved translation followed by either a two-step process or a terminology-oriented simultaneous extraction and normalization approach. Evaluation used French, English, and bilingual annotated datasets.
Main Results:
- The native French NLP method achieved a significantly higher overall F1-score of 0.51 (95% CI 0.47-0.55).
- Translated English methods yielded lower F1-scores of 0.39 (95% CI 0.34-0.44) and 0.38 (95% CI 0.36-0.40).
- Performance evaluation covered NER, normalization, and translation stages.
Conclusions:
- Native French NLP models demonstrate superior performance for extracting and normalizing French medical concepts compared to translated English approaches.
- Even with limited annotated French data, native French models are more effective for processing French medical texts.
- Improvements in translation models do not fully bridge the performance gap observed between native and translated methods.
Related Concept Videos
Improving Translational Accuracy
Leaky Scanning
Translation
Translation Produces the Building Blocks of Life
Proteins are...
Initiation of Translation
First, the initiator tRNA must be selected from the pool of elongator tRNAs by eukaryotic initiation factor 2 (eIF2). The initiator tRNA (Met-tRNAi) has conserved sequence elements including modified bases at...
Termination of Translation
Proteins: From Genes to Degradation
Transcription is the synthesis of RNA...

