Related Experiment Video
Updated: Aug 13, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
A unified framework and benchmark for generalizable biomedical knowledge extraction and applications with large
Wuyang Lan1, Siqi Zhang2, Wenzheng Wang1
1State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, Beijing, China; School of Information and Communication Engineering, Beijing University of Posts and Telecommunications, Beijing, China.
InfoFlowEX enhances biomedical information extraction (BIE) by unifying diverse datasets and employing novel tuning strategies. This framework improves large language model (LLM) adaptability for real-world biomedical knowledge discovery.
Area of Science:
- Biomedical Informatics
- Natural Language Processing
- Artificial Intelligence
Background:
- Biomedical information extraction (BIE) is crucial for converting unstructured text into computable knowledge.
- Current large language models (LLMs) face limitations due to heterogeneous datasets and a lack of unified benchmarks in BIE.
- Existing methods struggle with dataset variability, hindering generalizable knowledge extraction.
Purpose of the Study:
- To introduce InfoFlowEX, a unified framework designed for generalizable BIE using LLMs.
- To address the challenges of dataset heterogeneity and benchmark limitations in biomedical knowledge extraction.
- To enhance the adaptability and performance of LLMs in diverse biomedical applications.
Main Methods:
- Developed an automated data integration pipeline with ontology-guided alignment to create BIE-Corpus, a benchmark of 40 public datasets.
- Introduced a task-conditioned schema instruction tuning strategy to encode 28 biomedical entity and relation types.
- Enabled LLMs to align heterogeneous annotations and generalize across different BIE tasks.
Main Results:
- InfoFlowEX demonstrated robust adaptability and consistent performance gains over baseline methods.
- The framework achieved significant improvements with minimal task-specific customization.
- Evaluated successfully in evidence retrieval, clinical diagnosis from electronic health records, and knowledge graph expansion.
Conclusions:
- InfoFlowEX provides a unified and generalizable approach to biomedical knowledge extraction with LLMs.
- The framework effectively handles dataset heterogeneity and improves LLM performance across various applications.
- InfoFlowEX is highlighted as a valuable tool for real-world biomedical data analysis and knowledge discovery.
