Evaluating the Accuracy of ChatGPT-4o in Addressing Complex Clinical Questions Based on NCCN Guidelines for Rectal

Ryan Meyer1, Tamir E Bresler1, Kevin M Palmer1

  • 1Department of Surgery, Los Robles Regional Medical Center, Thousand Oaks, California, USA.

Abstract

Insights

Large language models like ChatGPT-4o show high accuracy in following rectal cancer guidelines, demonstrating improved performance over previous versions. However, complex treatment decisions remain a challenge for these AI tools.

Area of Science:

  • Oncology
  • Artificial Intelligence
  • Medical Guidelines

Background:

  • Rectal adenocarcinoma management involves complex guidelines with variable adherence.
  • Current large language models' (LLMs) ability to navigate these guidelines is not well-understood.

Purpose of the Study:

  • To evaluate the accuracy of ChatGPT-4o in adhering to NCCN Rectal Cancer Guidelines.
  • To assess LLM performance across different clinical domains within rectal cancer management.

Main Methods:

  • A cross-sectional study using 135 clinical questions derived from NCCN guidelines.
  • ChatGPT-4o was queried, and responses were rated by physician reviewers.
  • Inter-rater reliability and performance across domains were analyzed.

Main Results:

  • ChatGPT-4o achieved 94.1% Correct and 89.6% Accurate responses.
  • Performance was consistent across clinical domains, with perfect inter-rater agreement.
  • Errors were concentrated in complex, multi-step treatment decision points.

Conclusions:

  • ChatGPT-4o demonstrates high concordance with NCCN rectal cancer guidelines.
  • LLMs excel at factual recall but may struggle with complex clinical reasoning.
  • Further validation is needed before clinical integration.

Related Concept Videos