Related Experiment Video
Updated: Jan 10, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Benchmarking retrieval-augmented large language models in biomedical NLP: Application, robustness, and self-awareness
Mingchen Li1, Zaifu Zhan2, Han Yang3
1Division of Computational Health Sciences, Department of Surgery University of Minnesota, Minneapolis, MN, USA.
Retrieval-augmented large language models (LLMs) show promise in biomedical natural language processing (NLP) tasks. However, they require further development for improved robustness and self-awareness in complex scenarios.
Area of Science:
- Biomedical Natural Language Processing (NLP)
- Artificial Intelligence (AI)
- Machine Learning (ML)
Background:
- Large language models (LLMs) can generate hallucinations.
- Retrieval-augmented LLMs (RALs) mitigate hallucinations by retrieving external knowledge.
- The efficacy of RALs in biomedical NLP tasks is not well-established.
Purpose of the Study:
- To introduce a comprehensive benchmark for evaluating RALs in biomedical NLP.
- To assess RALs' performance across various tasks and robustness testbeds.
- To propose methods for enhancing RALs' robustness and negative awareness.
Main Methods:
- Developed the Biomedical Retrieval-Augmented Generation Benchmark (BARGE).
- Evaluated RALs on five biomedical NLP tasks and 11 datasets.
- Utilized four testbeds: unlabeled, counterfactual, diverse robustness, and self-awareness.
- Proposed a detect-and-correct strategy and contrastive learning for improvement.
Main Results:
- RALs generally outperform standard LLMs in biomedical NLP.
- RALs exhibit limitations in robustness and self-awareness, especially in counterfactual and diverse scenarios.
- Proposed methods significantly enhance robustness to unlabeled and counterfactual data.
- Improved models' ability to detect and avoid incorrect predictions.
Conclusions:
- Current RALs show potential but require refinement for biomedical applications.
- Robustness and self-awareness remain critical challenges for RALs in healthcare.
- Further research is needed to ensure the reliability and accuracy of RALs in high-stakes biomedical settings.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
09:20Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy