Related Experiment Video
Updated: Sep 17, 2025

05:25
Step By Step: Microsurgical training method combining two nonliving animal models
Published on: May 9, 2015
15.4K
A structured evaluation of LLM-generated step-by-step instructions in cadaveric brachial plexus dissection
Fulya Temizsoy Korkmaz1, Fatma Ok2, Burak Karip2
1Department of Anatomy, Hamidiye Faculty of Medicine, University of Health Sciences, Istanbul, 34668, Türkiye. fulya.temizsoykorkmaz@sbu.edu.tr.
BMC Medical Education
|July 2, 2025
Summary
Large language models (LLMs) show promise as educational tools for cadaver dissection, with ChatGPT-4o and Grok 3.0 excelling in scientific accuracy and guidance. These AI systems offer potential for scalable, individualized anatomy training, especially in resource-limited settings.
Area of Science:
- Medical Education
- Artificial Intelligence in Medicine
- Anatomical Dissection
Background:
- Large language models (LLMs) are increasingly used in various medical fields.
- Their performance in sensory-motor interactive environments like cadaver dissection remains unevaluated.
- Brachial plexus dissection presents a complex anatomical challenge for AI guidance.
Purpose of the Study:
- To comparatively analyze LLM performance in cadaver dissection.
- Evaluate scientific quality, educational value, and readability of LLM responses.
- Assess LLMs as step-by-step guides in a complex anatomical environment.
Main Methods:
- Developed a 28-item question set for brachial plexus dissection.
- Experienced anatomists blindly evaluated LLM responses using modified DISCERN and Global Quality Score.
- Assessed readability with Flesch Reading Ease, Flesch-Kincaid Grade Level, SMOG, Gunning Fog, and Coleman-Liau indices.
- Validated content using Content Validity Index and inter-rater reliability via ICC and Cohen's Kappa.
Main Results:
- ChatGPT-4o and Grok 3.0 achieved the highest scores for scientific accuracy and guidance structure (p < 0.01).
- DeepSeek demonstrated high readability but lacked content depth; Gemini performed moderately.
- Readability metrics significantly correlated with overall quality scores.
Conclusions:
- LLMs can provide scalable, individualized support for anatomy training, complementing traditional methods.
- This study is a foundational reference for AI-assisted cadaveric studies and surgical anatomy decision support.
- LLMs show potential in resource-limited educational settings, addressing mentorship or cadaver availability challenges.
Keywords:
Anatomical variationArtificial intelligence(ai) in anatomyBrachial plexusCadaveric dissectionClinical decision supportDissection guidanceEducational AILarge language models (LLMs)Readability analysisSurgical anatomy
