Related Experiment Video
Updated: Jul 2, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Leveraging Large Language Models for Clinical Abbreviation Disambiguation.
Manda Hosseini1, Mandana Hosseini2, Reza Javidan2
1Department of Computer Engineering, Zand Institute of Higher Education, Shiraz, Iran. manda.hosseinii@gmail.com.
This study enhances clinical abbreviation disambiguation using Large Language Models (LLMs) and a generative model (BIOGPT) for data augmentation. The approach improves accuracy, especially for rare abbreviations, by generating contextually relevant examples.
Area of Science:
- Biomedical Informatics
- Natural Language Processing
- Clinical Data Analysis
Background:
- Clinical abbreviation disambiguation is essential for medical information retrieval.
- Existing methods face challenges with limited data and ambiguous abbreviations.
- Accurate interpretation of abbreviations is vital for clinical text analysis.
Purpose of the Study:
- To improve clinical abbreviation disambiguation using Large Language Models (LLMs) and data augmentation.
- To address challenges of limited instances and ambiguous interpretations in clinical texts.
- To leverage BIOGPT for generating contextually relevant data to enhance disambiguation accuracy.
Main Methods:
- Utilized BlueBERT and Transformers for contextual understanding within LLMs.
- Employed a Generative Model (BIOGPT) pretrained on biomedical literature for data augmentation.
- Generated diverse, contextually relevant instances of clinical text for abbreviation expansion.
- Evaluated the approach on the CASI dataset, partitioned into training, validation, and test sets.
Main Results:
- Data augmentation with the Generative Model significantly improved disambiguation performance.
- Enhanced performance was particularly notable for abbreviation senses with limited instances.
- The method effectively addressed dataset imbalance and challenges from similar clinical concepts.
- The proposed model achieved good accuracy on the test set, outperforming previous methods.
Conclusions:
- LLMs and generative techniques are effective for clinical abbreviation disambiguation.
- Data augmentation using BIOGPT successfully addresses data scarcity and ambiguity.
- The approach offers a robust solution for improving the accuracy of clinical text analysis.
More Related Videos
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Related Concept Videos
Guidelines for Nursing Documentation I
Factual:
The following points emphasize the significance of upholding accurate and unbiased documentation in healthcare.
Language and Cognition
Improving Translational Accuracy
Leaky Scanning