Related Experiment Video
Updated: Nov 24, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
UMLS-based data augmentation for natural language processing of clinical research literature
Tian Kang1, Adler Perotte1, Youlan Tang1
1Department of Biomedical Informatics, Columbia University, New York, New York, USA.
We developed UMLS-EDA, a knowledge-based data augmentation method, to enhance deep learning models in biomedical natural language processing. This approach significantly improves named entity recognition and sentence classification performance, especially in low-resource settings.
Area of Science:
- Biomedical Natural Language Processing
- Machine Learning
- Data Science
Background:
- Deep learning models in biomedical NLP often suffer from limited training data.
- Existing data augmentation techniques may not fully leverage domain-specific knowledge.
- Improving model performance in low-resource biomedical domains is crucial for advancing research and clinical applications.
Purpose of the Study:
- To develop and evaluate a novel knowledge-based data augmentation method for biomedical NLP.
- To enhance the performance of deep learning models by addressing training data scarcity.
- To investigate the effectiveness of incorporating the Unified Medical Language System (UMLS) knowledge into data augmentation.
Main Methods:
- Extended the easy data augmentation (EDA) method by integrating UMLS knowledge, creating UMLS-EDA.
- Designed experiments to evaluate UMLS-EDA on deep learning architectures for Named Entity Recognition (NER) and classification tasks.
- Compared the performance of UMLS-EDA against baseline models and BERT-based approaches.
Main Results:
- UMLS-EDA significantly improved NER performance for LSTM-CRF models (micro-F1 scores: +5%, +17%, +15%).
- The LSTM-CRF model with UMLS-EDA outperformed BERT transfer learning for NER (0.66 vs. 0.63 micro-F1).
- UMLS-EDA enhanced sentence classification models, achieving a 9% micro-F1 score increase (0.75 to 0.84), surpassing BERT pretraining (0.82).
Conclusions:
- UMLS-EDA is an effective knowledge-based data augmentation method for biomedical NLP.
- The method substantially improves deep learning models for both NER and sentence classification.
- UMLS-EDA offers valuable insights for developing superior deep learning approaches in low-resource biomedical domains.
More Related Videos
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Related Concept Videos
Clinical Trials: Overview
Clinical Trials
There are four phases in a clinical trial. A phase one...
Improving Translational Accuracy
Upsampling