Related Experiment Video
Updated: Jan 13, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance of Seven Large Language Models on Anatomy Examination Questions
Weronika Chaba-Karnaś1, Natalia Kozioł1, Natalia Kalita1
1Department of Anatomy, Jagiellonian University Medical College, Kraków, Poland.
Abstract:
Artificial intelligence is among the most rapidly developing branches of technology. It has proven to be a helpful tool in various fields, including medicine. Significant advances in the development of new language models prompt an evaluation of their effectiveness across various areas of medicine, including anatomy. This study aimed to assess the effectiveness of artificial intelligence in solving theoretical anatomy exams designed for medical students. The study utilized 555 multiple-choice questions (150 in Polish and 405 in English) sourced from past anatomy exams for the medical program. The models tested included: ChatGPT-4o mini, ChatGPT-4o, DeepSeek, Copilot, Gemini, and two Polish models: Bielik and PLLum. Each question was asked only once. For analysis purposes, the questions were categorized by type and by the anatomical structure they addressed. Out of 555 questions, ChatGPT-4o mini answered 394 correctly (71%), ChatGPT-4o - 461 (83.1%), DeepSeek - 427 (76.9%), Copilot - 442 (79.6%), Gemini - 439 (78.8%), Bielik - 166 (29.9%), and PLLum - 222 (40.0%). The language models performed poorest on multiple-answer questions (37.6%) and best on questions concerning the function of a given organ (75%). Most of the tested language models are capable of independently passing the exam, which should serve as a warning to teaching staff supervising students during exams and assessments. Properly formulated questions can currently hinder students relying on artificial intelligence from passing, but ongoing AI advancements may result in even higher pass rates in the future.
More Related Videos
05:56Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Related Concept Videos
Classification of Bones
Long and Short Bones
The appendicular skeleton, particularly the upper and lower limbs, is primarily made of long and short bones. The...
Improving Translational Accuracy
Improving Translational Accuracy