Related Experiment Video
Updated: Apr 12, 2026

08:11
Surgical Training for the Implantation of Neocortical Microelectrode Arrays Using a Formaldehyde-fixed Human Cadaver Model
Published on: November 19, 2017
11.3K
Using large language models (ChatGPT, Copilot, PaLM, Bard, and Gemini) in Gross Anatomy course: Comparative analysis
Volodymyr Mavrych1, Paul Ganguly1, Olena Bolgova1
1College of Medicine, Alfaisal University, Riyadh, Kingdom of Saudi Arabia.
Summary
Generative artificial intelligence large language models (LLMs) show varied accuracy in medical education. ChatGPT-4 demonstrated superior performance in answering medical multiple-choice questions and generating clinical scenarios compared to other tested LLMs.
Area of Science:
- Medical Education
- Artificial Intelligence
- Anatomy
Background:
- Generative artificial intelligence large language models (LLMs) are increasingly used in various fields, including medical education.
- Concerns exist regarding the accuracy and reliability of LLMs in educational contexts.
- Assessing LLM performance in specialized subjects like Gross Anatomy is crucial for understanding their potential role.
Purpose of the Study:
- To compare the accuracy and proficiency of six different large language models (LLMs) in medical education.
- To evaluate LLMs' ability to answer medical multiple-choice questions (MCQs) and generate clinical scenarios for Gross Anatomy.
- To determine the best-performing LLM for tasks relevant to medical student education in Gross Anatomy.
Main Methods:
- Six LLMs (ChatGPT-4, ChatGPT-3.5-turbo, ChatGPT-3.5, Copilot, PaLM, Bard, and Gemini) were tested.
- LLMs answered 50 USMLE-style Gross Anatomy MCQs, with results evaluated over five attempts for accuracy, relevance, and comprehensiveness.
- LLMs generated clinical scenarios and MCQs for specific upper limb topics, with expert grading on a 0-5 scale.
Main Results:
- ChatGPT-4 achieved the highest accuracy (60.5% ± 1.9%) in answering MCQs, significantly outperforming other LLMs (p < 0.05).
- In generating clinical scenarios and MCQs, ChatGPT-4 also yielded the best results, followed by Gemini and ChatGPT-3.5 variants.
- Copilot, Google PaLM 2, and Bard showed lower performance in both MCQ answering and scenario generation tasks.
Conclusions:
- While LLMs show promise as supplementary tools, they are not yet mature enough to replace human educators in Gross Anatomy.
- ChatGPT-4 demonstrated the most robust performance among the evaluated LLMs for medical education tasks.
- Further development is needed for LLMs to fully support medical teaching and learning effectively and reliably.

