Related Experiment Video
Updated: Sep 5, 2026

Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025
Prompt Strategy and Model Choice in AI-Assisted Gross Anatomy Education: Evaluating Zero-Shot, Few-Shot, and
Volodymyr Mavrych1, Hassan S Shaibah1, Akef Obeidat1
1College of Medicine, Alfaisal University, Saudi Arabia.
Abstract:
Visual identification of anatomical structures is a foundational skill in gross anatomy education. Whether multimodal large language models (LLMs) can reliably perform this task, and whether prompt engineering can meaningfully improve their accuracy, remains insufficiently investigated This cross-sectional comparative study evaluated four leading multimodal LLMs (GPT-5.2, Claude Sonnet 4.6, Gemini 3, and Grok 4) on their ability to identify anatomical structures across cadaveric dissection images. A total of 75 expert-validated anatomical structures spanning five body regions (abdomen, head and neck, lower limb, thorax, and upper limb) were evaluated using three prompt strategies: zero-shot (P0), few-shot (PF), and chain-of-thought (PCoT). Each prompt strategy was administered across three independent sessions conducted in April-May 2026, yielding 675 binary-scored responses per model (2700 total responses). Gemini 3 achieved the highest overall accuracy (67.3%), followed by GPT-5.2 (41.9%), Claude Sonnet 4.6 (27.3%), and Grok 4 (23.0%)-an ordering that inverts the hierarchy typically observed in text-based anatomy assessments, where GPT-4o has generally led, and Gemini has ranked lower. Gemini 3 significantly outperformed all other models (all Bonferroni-corrected p < 0.001). Prompt strategy had a statistically significant effect for Gemini 3 (PCoT > PF, p = 0.015) and Grok 4 (PCoT > P0, p = 0.023); no significant prompt effect was observed for GPT-5.2 or Claude Sonnet 4.6. Performance varied substantially by anatomical region: the abdomen consistently yielded the highest accuracy across all models, while the lower limb yielded the lowest. Inter-trial consistency was high (70.7%-89.3%) but dissociated from accuracy, as some models produced reproducibly incorrect responses. Current multimodal LLMs, accessed via consumer interfaces, are insufficient for reliable standalone identification of cadaveric anatomical structures. Gemini 3, when used with chain-of-thought prompting, may serve as a supplementary aid in select anatomical regions; however, critical educator supervision and verification against authoritative anatomical resources remain essential before any clinical or educational deployment.