Related Experiment Video
Updated: Aug 4, 2026

10:42
Rodent Behavioral Testing to Assess Functional Deficits Caused by Microelectrode Implantation in the Rat Motor Cortex
Published on: August 18, 2018
8.8K
Matching Human Expertise: ChatGPT's Performance on Hand Surgery Examinations
Zachary A Kirschenbaum1, Yuri Han1, Kiera L Vrindten1
1Rutgers Robert Wood Johnson Medical School, New Brunswick, NJ, USA.
Summary
Artificial intelligence (AI) tools like ChatGPT 4o show human-equivalent performance on hand surgery exams, especially with improved prompts and literature access. This demonstrates AI's potential to aid medical education and practice.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Education Technology
- Surgical Assessment Tools
Background:
- Artificial intelligence (AI) integration in healthcare is rapidly advancing, with tools like ChatGPT showing promise.
- Previous evaluations indicated that ChatGPT 3.5 underperformed compared to human experts on hand surgery self-assessment exams.
- The study addresses the need to assess newer AI capabilities in specialized medical fields.
Purpose of the Study:
- To evaluate the performance of ChatGPT 4o on American Society for Surgery of the Hand (ASSH) self-assessment questions.
- To determine if enhanced techniques, such as improved prompts and file search, can increase AI accuracy.
- To compare ChatGPT 4o's performance against previous versions and human benchmarks.
Main Methods:
- Utilized data from ASSH self-assessment examinations (2008-2013).
- Investigated the impact of ChatGPT model version, prompt engineering, and file search on response accuracy via OpenAI's API.
- Performed statistical analysis using one-way ANOVA and assessed test reliability with KR-20.
Main Results:
- ChatGPT 4o, particularly with enhanced prompting and literature access, achieved human-comparable performance on text-based questions.
- ChatGPT 4o significantly outperformed ChatGPT 3.5, with notable improvements from advanced prompting and file search.
- The 2013 examination demonstrated high reliability with a KR-20 score of 0.946.
Conclusions:
- AI, specifically ChatGPT 4o, can perform at a human-equivalent level on specialized hand surgery self-assessment examinations.
- Findings suggest AI's potential as a valuable supplementary tool in medical education.
- AI tools may offer supportive resources for clinical practice and professional development.

