Related Experiment Video
Updated: Jun 30, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
554
Ensemble pretrained language models to extract biomedical knowledge from literature.
Zhao Li1, Qiang Wei1, Liang-Chin Huang1
1McWilliams School of Biomedical Informatics, University of Texas Health Science Center at Houston, Houston, TX 77030, United States.
Summary
A novel Natural Language Processing (NLP) system achieved top rankings in biomedical named entity recognition (NER) and relation extraction (RE), outperforming large language models. Task-specific models demonstrate superior performance for biomedical text mining.
Area of Science:
- Biomedical Informatics
- Computational Linguistics
- Natural Language Processing (NLP)
Background:
- The growing volume of biomedical literature requires automated methods for extracting relationships between concepts.
- Developing robust NLP techniques is crucial for building knowledge bases and identifying research gaps.
Purpose of the Study:
- To develop and evaluate an NLP system for the LitCoin challenge, focusing on named entity recognition (NER) and relation extraction (RE).
- To benchmark NLP methodologies using a manually annotated corpus.
Main Methods:
- Ensemble learning combining BioBERT, PubMedBERT, and BioM-ELECTRA for NER.
- A rule-driven method for cell line and taxonomy detection.
- Finetuning the T0pp model for relation extraction, incorporating entity location information.
Main Results:
- Achieved first place in NER and second place in relation extraction and novelty prediction in the LitCoin challenge.
- The finetuned model significantly outperformed general-purpose large language models like ChatGPT 3.5 and 4 in a zero-shot setting.
Conclusions:
- The developed NLP system demonstrates high efficacy in NER and RE for biomedical entities.
- Task-specific NLP models offer superior performance compared to generic large language models for specialized biomedical text analysis.
Keywords:
ensemble learningknowledge baselarge language modelnamed entity recognitionrelation extraction
