Related Experiment Video
Updated: Aug 13, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Intelligent Framework for Adverse Drug Event Identification Using Large Language Models and Retrieval-Augmented
Junlong Ma1, Xuehong Wu2, Zeying Feng3
1Department of Pharmacy, Xiangya Hospital, Central South University, Changsha, Hunan, China, 1 0731 88618933.
Journal of Medical Internet Research
|August 10, 2026
Summary
Retrieval-augmented generation (RAG) significantly improves adverse drug event (ADE) identification in Chinese clinical notes using large language models (LLMs). This RAG approach enhances accuracy and recall, offering a robust framework for pharmacovigilance and drug safety.
Area of Science:
- Natural Language Processing
- Artificial Intelligence in Medicine
- Pharmacovigilance
Background:
- Adverse drug events (ADEs) present significant public health and economic challenges.
- Extracting ADE information from unstructured clinical notes is difficult due to semantic complexity.
- Large language models (LLMs) show promise for text comprehension but can suffer from domain-specific hallucinations.
Purpose of the Study:
- To evaluate the effectiveness of retrieval-augmented generation (RAG) for identifying ADEs in Chinese clinical narratives using LLMs.
- To establish a paradigm for ADE identification in clinical text using RAG and LLMs.
Main Methods:
- A dataset of 18,432 high-quality Chinese clinical notes was curated and annotated.
- A gold-standard reference dataset (n=2510) and an ADE knowledge base (n=5144) were established.
- Three LLMs (DeepSeek-V3, ERNIE 3.5-8K, GPT-4o) were evaluated using nonaugmented generation (NAG), static-augmented generation (SAG), and RAG strategies, with performance assessed via precision, recall, and F1-score at multiple recognition levels.
Main Results:
- The first Chinese ADE corpus from clinical notes was created and released.
- RAG consistently outperformed NAG and SAG, achieving an optimal L3 F1-score of 0.9638 with DeepSeek-V3.
- RAG significantly improved GPT-4o's recall from 0.6419 to 0.9241 and demonstrated clinical utility with high specificity (0.9821) on real-world data.
Conclusions:
- Integrating a domain-specific knowledge base with LLMs via RAG is effective for accurate ADE identification in Chinese clinical notes.
- This RAG strategy mitigates LLM hallucinations, providing a benchmark for pharmacovigilance and drug safety research.
- The developed framework supports clinical decision support systems and advances drug safety monitoring.
