Performance of Publicly Available Large Language Models on Internal Medicine Board-style Questions

Constantine Tarabanis1, Sohail Zahid1, Marios Mamalis2

  • 1Leon H. Charney Division of Cardiology, NYU Langone Health, New York University School of Medicine, New York, New York, United States of America.

PLOS Digital Health
|September 17, 2024
PubMed
Summary

Large language models (LLMs) show promise in medical exams, with GPT-4.0 outperforming human physicians on internal medicine board questions. Augmenting LLMs with medical texts improves their accuracy.