Related Experiment Video
Updated: Jan 7, 2026

Grossing of Non-neoplastic Globes, Including Fetal Eyes
Published on: May 30, 2025
ChatGPT's performance on a specialist forensic pathology examination: implications for forensic pathologists and
Hans H de Boer1,2, Gregory Young3,4, Heinrich Bouwer3,4
1Department of Forensic Medicine, Monash University, 65 Kavanagh Street, Southbank, VIC, 3006, Australia. hans.de.boer@vifm.org.
Background:
The use of artificially intelligent Language Models (LLMs) like ChatGPT is increasing rapidly in medicine, but their accuracy, reasoning quality, and contextual safety remains an issue. This is especially relevant for forensic pathology, where balanced reasoning, contextual sensitivity, and precise communication are essential. We performed an in-depth assessment of ChatGPT's capabilities and limitations for forensic pathology, which also improves our understanding of the risks and benefits of LLM use in medicine more broadly.
Methods:
ChatGPT-4.5 Turbo's performance was tested using a multifaceted mock exam, consisting of core elements of the forensic pathology fellowship exam of the Royal College of Pathologists of Australasia (RCPA). The mock exam included essay-style questions, image-based tasks, and case reporting. ChatGPT's responses were blindly marked by experienced examiners according to standard RCPA criteria, assessing factual accuracy, reasoning structure, and communication quality.
Results:
ChatGPT performed well on most essay-style knowledge questions, achieving higher scores on topics with well-established knowledge. Performance was however poor for tasks requiring complex reasoning, image interpretation, or the context-dependent analysis of autopsy findings. Importantly, ChatGPT's output was always phrased fluently and persuasively, creating an impression of confidence that was independent of factual accuracy.
Conclusions:
ChatGPT can reliably reproduce well-established forensic pathology knowledge. However, it lacks the capabilities needed for higher-level tasks and often generates unjustifiably confident and misleading output. Its use may be acceptable for low-risk administrative or educational purposes, provided output is carefully reviewed by qualified experts. Non-specialists should not rely on such tools for forensic pathology information. Continued evaluation is needed to ensure safe and responsible use.

