Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Improving Translational Accuracy02:07

Improving Translational Accuracy

14.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.0K
Improving Translational Accuracy02:07

Improving Translational Accuracy

3.5K
3.5K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

RiTeK: A Dataset for Large Language Models Complex Reasoning over Textual Knowledge Graphs in Medicine.

Findings of ACL. ACL·2026
Same author

Benchmarking information extraction of physical activity from electronic health record with large language models: an natural language processing pipeline and comparative evaluation.

Journal of the American Medical Informatics Association : JAMIA·2026
Same author

Let large language models judge each other: multi-agent peer-reviewed reasoning for medical question answering.

Journal of the American Medical Informatics Association : JAMIA·2026
Same author

Antidiabetic Drug Associations With Heart Failure Outcomes: Real-World Evidence Study Using Electronic Health Records.

JMIR diabetes·2026
Same author

Retrieval-augmented in-context learning for multimodal large language models in disease classification.

Journal of biomedical informatics·2026
Same author

NutriRAG: unleashing the power of large language models for food identification and classification through retrieval methods.

Journal of the American Medical Informatics Association : JAMIA·2026

Related Experiment Video

Updated: Jan 10, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1000

Benchmarking retrieval-augmented large language models in biomedical NLP: Application, robustness, and

Mingchen Li1, Zaifu Zhan2, Han Yang3

  • 1Division of Computational Health Sciences, Department of Surgery University of Minnesota, Minneapolis, MN, USA.

Science Advances
|November 21, 2025
PubMed
Summary

Retrieval-augmented large language models (LLMs) show promise in biomedical natural language processing (NLP) tasks. However, they require further development for improved robustness and self-awareness in complex scenarios.

More Related Videos

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

1.3K
Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
09:20

Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications

Published on: February 23, 2019

9.2K

Related Experiment Videos

Last Updated: Jan 10, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1000
Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

1.3K
Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
09:20

Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications

Published on: February 23, 2019

9.2K

Area of Science:

  • Biomedical Natural Language Processing (NLP)
  • Artificial Intelligence (AI)
  • Machine Learning (ML)

Background:

  • Large language models (LLMs) can generate hallucinations.
  • Retrieval-augmented LLMs (RALs) mitigate hallucinations by retrieving external knowledge.
  • The efficacy of RALs in biomedical NLP tasks is not well-established.

Purpose of the Study:

  • To introduce a comprehensive benchmark for evaluating RALs in biomedical NLP.
  • To assess RALs' performance across various tasks and robustness testbeds.
  • To propose methods for enhancing RALs' robustness and negative awareness.

Main Methods:

  • Developed the Biomedical Retrieval-Augmented Generation Benchmark (BARGE).
  • Evaluated RALs on five biomedical NLP tasks and 11 datasets.
  • Utilized four testbeds: unlabeled, counterfactual, diverse robustness, and self-awareness.
  • Proposed a detect-and-correct strategy and contrastive learning for improvement.

Main Results:

  • RALs generally outperform standard LLMs in biomedical NLP.
  • RALs exhibit limitations in robustness and self-awareness, especially in counterfactual and diverse scenarios.
  • Proposed methods significantly enhance robustness to unlabeled and counterfactual data.
  • Improved models' ability to detect and avoid incorrect predictions.

Conclusions:

  • Current RALs show potential but require refinement for biomedical applications.
  • Robustness and self-awareness remain critical challenges for RALs in healthcare.
  • Further research is needed to ensure the reliability and accuracy of RALs in high-stakes biomedical settings.