Related Experiment Video
Updated: Nov 21, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Medical concept normalization in French using multilingual terminologies and contextual embeddings
Perceval Wajsbürt1, Arnaud Sarfati2, Xavier Tannier1
1Sorbonne Université, Inserm, Université Sorbonne Paris Nord, Laboratoire d'Informatique Médicale et d'Ingénierie des Connaissances pour la e-Santé, LIMICS, 75006 Paris, France.
This study introduces a novel system for French concept normalization, improving medical text analysis without translation. The approach achieves strong results, outperforming existing methods even without labeled data.
Area of Science:
- Natural Language Processing
- Medical Informatics
- Computational Linguistics
Background:
- Concept normalization links medical terms to standardized terminologies like UMLS®.
- Traditional methods struggle with non-English languages due to resource limitations.
- Developing effective concept normalization for diverse languages is crucial for global medical informatics.
Purpose of the Study:
- To present a novel system for concept normalization in French medical texts.
- To leverage multilingual terminologies and embedding models for improved accuracy.
- To achieve effective concept normalization without direct translation or supervision.
Main Methods:
- The system treats concept normalization as a highly-multiclass classification problem.
- Contextualized embeddings encode medical terms, classified using cosine similarity and softmax.
- A two-step training process involves finetuning embeddings and further training with hard negative selection.
Main Results:
- The proposed approach achieves good results on French medical corpora, even with no labeled data.
- Performance surpasses existing supervised methods that utilize labeled data.
- Training with both French and English terms significantly boosts performance on French benchmarks.
Conclusions:
- The developed distantly supervised method is applicable across various medical domains and document types.
- This approach offers a simpler and more effective multilingual strategy for processing non-English medical texts.
- The findings pave the way for broader advancements in multilingual medical text analysis.
More Related Videos
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
09:20Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
Related Concept Videos
Natural and Artificial Concepts
Anatomical Terminology
Empathy
Lateralization
Language and Cognition
Regional Terms
Primarily, the human body has two major regions, the axial and appendicular regions. The axial region comprises regions from the head to the abdomen and makes up the central body axis. In contrast,...