Related Experiment Video
Updated: May 20, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
The performance of ChatGPT and other large language models on multiple-choice questions in biomedical disciplines: A
Colleen M Cheverko1, Volodymyr Mavrych2, Olena Bolgova2
1Department of Anatomy and Cell Biology, Rush University, Chicago, Illinois, USA.
Abstract:
While large language models (LLMs) have shown promise as learning tools for medical education, their reported accuracy on multiple-choice questions (MCQs) varies widely across studies, necessitating synthesis. This meta-analysis synthesizes LLM accuracy on text-based MCQs from biomedical disciplines and USMLE Step 1-level content and explores study characteristics that moderate LLM performance. Studies published between January 1, 2022, and August 5, 2025, were identified using seven databases. Titles and abstracts were screened against eligibility criteria, which included testing LLM performance on text-based MCQs in English. Extracted data included accuracy rates (correct responses/total questions) and study characteristics. A random-effects proportional meta-analysis was used to calculate the pooled accuracy of LLM versions across studies, and a meta-regression tested the effects of study characteristics on accuracy. The final analysis included 41 articles comprising 207 proportions. Newer OpenAI models (GPT-4o through GPT o1 = 90%; GPT-4 = 82%) demonstrated significantly higher accuracy than earlier versions (GPT-3/3.5 = 59%; p < 0.001). Model version explained 56% of the total variance (p < 0.001). The pooled accuracy of newer Anthropic (Claude 3 through 3.7 = 86%) and DeepSeek (R1 = 86%) versions fell within the range of performance for the newer OpenAI versions, while newer Google (Gemini series = 76%) and Microsoft (Copilot = 69%) versions demonstrated lower performance. A discipline-level analysis demonstrated that OpenAI's GPT-4 through GPT-o1 models achieved pooled accuracy rates between 83 and 87% for most disciplines, including the anatomical sciences. These findings demonstrate that newer LLMs are achieving higher accuracy than older models on biomedical and Step 1-level MCQs, supporting their potential integration as supplementary study tools in foundational biomedical education.
