Related Experiment Video
Updated: Dec 17, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
The Impact of Specialized Corpora for Word Embeddings in Natural Langage Understanding
Antoine Neuraz1,2, Bastien Rance1, Nicolas Garcelon1
1INSERM, UMR 1138 Team 22, Paris Descartes, Paris, France.
Abstract:
Recent studies in the biomedical domain suggest that learning statistical word representations (static or contextualized word embeddings) on large corpora of specialized data improve the results on downstream natural language processing (NLP) tasks. In this paper, we explore the impact of the data source of word representations on a natural language understanding task. We compared embeddings learned with Fasttext (static embedding) and ELMo (contextualized embedding) representations, learned either on the general domain (Wikipedia) or on specialized data (electronic health records, EHR). The best results were obtained with ELMo representations learned on EHR data for the two sub-tasks (+7% and +4% of gain in F1-score). Moreover, ELMo representations were trained with only a fraction of the data used for Fasttext.
Related Concept Videos
Natural and Artificial Concepts
Language and Cognition
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Components of Language
Improving Translational Accuracy
Improving Translational Accuracy

