Related Experiment Video
Updated: Jan 10, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Benchmarking retrieval-augmented large language models in biomedical NLP: Application, robustness, and self-awareness
Mingchen Li1, Zaifu Zhan2, Han Yang3
1Division of Computational Health Sciences, Department of Surgery University of Minnesota, Minneapolis, MN, USA.
Abstract:
To reduce hallucinations in large language models (LLMs), retrieval-augmented LLMs (RALs) retrieve supporting knowledge from external databases. However, their performance on biomedical natural language processing (NLP) tasks remains underexplored. We introduce Biomedical Retrieval-Augmented Generation Benchmark, a comprehensive evaluation framework assessing RALs across five biomedical NLP tasks and 11 datasets, using four testbeds: unlabeled robustness, counterfactual robustness, diverse robustness, and self-awareness. To improve RALs' robustness and negative awareness, we propose a detect-and-correct strategy and a contrastive learning approach. Experimental results show that RALs generally outperform standard LLMs on most biomedical tasks, but still struggle with robustness and self-awareness, particularly under counterfactual and diverse scenarios. Our proposed methods significantly improve performance in robustness to unlabeled and counterfactual data, and increase the model's ability to detect and avoid incorrect predictions. These findings highlight key limitations in current RALs and underscore the need for continued refinement to ensure reliability and accuracy in high-stakes biomedical applications.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
09:20Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy