Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Video

Updated: Jun 19, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

BRIDGE: benchmarking large language models for understanding real-world clinical practice texts.

Jiageng Wu1, Bowen Gu1, Ren Zhou2

  • 1Division of Pharmacoepidemiology and Pharmacoeconomics, Department of Medicine, Brigham and Women's Hospital, Harvard Medical School, Boston, MA, USA.

Nature Biomedical Engineering
|June 17, 2026
PubMed

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

SmartAlert - Implementing Machine Learning-Driven Clinical Decision Support for Inpatient Laboratory Utilization Reduction.

NEJM AI·2026
Same author

Why and How to Monitor Deployed AI Systems in Health Care.

NEJM catalyst innovations in care delivery·2026
Same author

The Comparative Effectiveness of Carvedilol Versus Other Nonselective β-Blockers in Cirrhosis.

Annals of internal medicine·2026
Same author

Bleeding Risk With Apixaban Versus Rivaroxaban: A Reference Trial Emulation Predicting the Results of COBRRA-VTE and COBRRA-AF Using US Health Care Claims.

Circulation. Population health and outcomes·2026
Same author

Risk of Hyperkalemia in Patients with Heart Failure Treated with Spironolactone in Combination with Sacubitril/Valsartan vs. Renin-Angiotensin System Inhibitors.

Clinical pharmacology and therapeutics·2026
Same author

Association between Interictal Spike Rate and Seizure Frequency in a Large Epilepsy Cohort.

medRxiv : the preprint server for health sciences·2026
Summary

A new benchmark, BRIDGE, evaluates large language models (LLMs) on real-world clinical data. It shows performance varies, with open-source LLMs matching proprietary ones and updated general models outperforming older specialized ones.

Area of Science:

  • Artificial Intelligence
  • Medical Informatics
  • Natural Language Processing

Background:

  • Large language models (LLMs) show potential in medicine, but real-world clinical data benchmarking is lacking.
  • Existing benchmarks often use artificial questions or limited datasets, not reflecting clinical practice complexity.
  • There's a need for comprehensive evaluation of LLMs on diverse, real-world clinical text.

Purpose of the Study:

  • To introduce BRIDGE, a multilingual benchmark for evaluating LLMs on real-world clinical data.
  • To assess the performance of various LLMs across different languages, tasks, and clinical specialties.
  • To provide a resource for advancing LLM development in clinical text understanding.

Main Methods:

  • Developed BRIDGE, a benchmark with 87 tasks from 59 real-world clinical data sources in 9 languages.

More Related Videos

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

Related Experiment Videos

Last Updated: Jun 19, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

  • Included 8 task types (e.g., triage, diagnosis, coding) across 14 clinical specialties.
  • Evaluated 95 LLMs using multiple inference strategies.
  • Main Results:

    • Significant performance variations were observed across LLM sizes, languages, tasks, and specialties.
    • Open-source LLMs demonstrated performance comparable to proprietary models.
    • Medically fine-tuned models on older architectures underperformed updated general-purpose LLMs.

    Conclusions:

    • BRIDGE offers a robust evaluation framework for LLMs in clinical settings.
    • The benchmark highlights the need for continuous development and evaluation of LLMs for real-world medical applications.
    • BRIDGE and its leaderboard serve as crucial resources for the research community.