Related Experiment Video
Updated: Sep 2, 2026

Dissection, MicroCT Scanning and Morphometric Analyses of the Baculum
Published on: March 19, 2017
AnatomyGPT 5.2 Versus ChatGPT 5.2 in Solving Anatomical Questions: A Comparative Study
Weronika Chaba-Karnaś1, Eliza Tatarczyk1, Natalia Kozioł1
1Department of Anatomy, Jagiellonian University Medical College, Kraków, Poland.
Abstract:
Artificial intelligence is increasingly being used in education, including the field of anatomy, where it could support learning. The aim of this study was to evaluate and compare the performance of two large language models (LLMs), AnatomyGPT 5.2 and ChatGPT 5.2, in answering anatomy-based questions. A standardized prompt was prepared, and anatomical questions, originally used in official anatomy tests, were posed separately to AnatomyGPT 5.2 and ChatGPT 5.2 in two independent attempts. A total of 550 questions were used, including 150 in Polish and 400 in English. ChatGPT 5.2 outperformed AnatomyGPT 5.2 in both trials, achieving accuracies of 84.36% and 86.36%, compared with 82.73% and 85.27%, respectively; however, the differences between the models were not statistically significant in either trial. In both trials, ChatGPT 5.2 and AnatomyGPT 5.2 performed better on the English question set than on the Polish question set, with the differences between the question sets being statistically significant (p < 0.05). Questions on innervation and vascularization were most frequently answered correctly by both models, while multiple-choice questions were least frequently correct; no statistically significant differences were found between models or trials across all categories of questions (p > 0.05). Cohen's kappa analysis indicated substantial agreement between repeated responses for both models across the two trials (p < 0.001). The intraclass correlation coefficient was 0.62 for AnatomyGPT 5.2 and 0.66 for ChatGPT 5.2, indicating moderate agreement between the original and repeated responses for both models. Overall, the studied models demonstrated generally high performance across most questions and relatively consistent responses over time. Higher accuracy was observed for the English question set than for the Polish question set; however, this difference cannot be attributed solely to language.

