Related Experiment Video
Updated: Jun 15, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Recognition and normalization of multilingual symptom entities using in-domain-adapted BERT models and classification
Fernando Gallego1,2, Francisco J Veredas1,2
1Departamento de Lenguajes y Ciencias de la Computación, Universidad de Málaga, Blvr. Louis Pasteur, 35, Puerto de la Torre, Málaga 29071, Spain.
This study advances clinical natural language processing for low-resource languages by developing novel methods for detecting and linking clinical entities. The new approach sets a state-of-the-art benchmark in multilingual clinical text analysis.
Area of Science:
- Biomedical Informatics
- Computational Linguistics
- Natural Language Processing
Background:
- Clinical natural language processing (NLP) is challenging due to limited annotated data, especially for low-resource languages.
- Developing robust methods for identifying and normalizing clinical entities (symptoms, signs, findings) in multilingual medical texts is crucial.
Purpose of the Study:
- To address the challenges in clinical NLP for low-resource languages by contributing to the detection and normalization of clinical entities.
- To develop and evaluate models for named entity recognition (NER) and named entity linking (NEL) within the SympTEMIST shared task.
Main Methods:
- For NER (Subtask 1), a BERT-based model assembly pretrained on an oncology corpus was utilized for Spanish clinical texts.
- For NEL (Subtasks 2 and 3), a classification strategy employing a contrastive-learning-trained bi-encoder (SapBERT-like models) was developed.
- Multilingual capabilities were achieved by translating a Spanish knowledge base into other languages using machine translation tools.
Main Results:
- The proposed approach achieved state-of-the-art results across all three subtasks of the SympTEMIST challenge.
- Subtask 1 (Spanish NER) yielded precision=0.804, F1-score=0.748, and recall=0.699.
- Subtask 2 (Spanish NEL) showed performance gains up to 5.5% in top-1 accuracy with a WNT-softmax layer.
- Subtask 3 (Multilingual NEL) demonstrated significant improvements, with the multilingual bi-encoder outperforming other models in most languages, achieving 13% (Portuguese) and 13.26% (Swedish) gains in top-1 accuracy.
Conclusions:
- The developed methods significantly advance the state-of-the-art in clinical entity detection and normalization for multilingual clinical texts.
- The approach demonstrates the effectiveness of BERT-based models and contrastive learning for named entity linking in low-resource clinical NLP scenarios.
- This work provides a strong foundation for future research in cross-lingual clinical information extraction and supports the development of NLP tools for diverse linguistic contexts.
Related Concept Videos
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Language and Cognition
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Force Classification
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Classification of Systems-II

