Related Experiment Video
Updated: Jan 8, 2026

Robot-assisted Total Mesorectal Excision and Lateral Pelvic Lymph Node Dissection for Locally Advanced Middle-low Rectal Cancer
Published on: February 12, 2022
Artificial intelligence takes on the multidisciplinary committee: A single-center study for rectal cancer management
Muhammad U Khalid1, Charbel El-Kefraoui1, Alina Wang1
1Department of Surgery, St Paul's Hospital, University of British Columbia, Vancouver, British Columbia, Canada.
Background:
As rectal cancer management evolves, the multidisciplinary committee becomes increasingly important in integrating expertise to optimize patient outcomes. Current artificial intelligence large language models have demonstrated preliminary capacity to apply medical guidelines to specific patient scenarios. This study assesses the ability of these publicly available artificial intelligence large language models (Gemini, Grok, ChatGPT) to predict multidisciplinary committee recommendations for rectal cancer.
Methods:
Adult patients who presented to the multidisciplinary committee at a Canadian tertiary hospital with a new diagnosis of rectal adenocarcinoma before March 2025 were sequentially and retrospectively included in the study. Baseline demographic characteristics were recorded. Redacted patient vignettes were presented to each artificial intelligence large language models, and concordance between artificial intelligence large language models and multidisciplinary committee management recommendations was graded on a 5-point Likert scale by 3 independent reviewers. The Cohen κ coefficient was used to assess inter-rater agreement, and descriptive statistics, odds ratios, and multivariable regression used to assess each artificial intelligence large language model's performance.
Results:
One hundred patients were included, with a median age of 60 years (range, 38-90 years). Most patients were male (70%), with a mean Charlson comorbidity index of 4.37 (range, 2-10). All 4 stages of rectal cancer were represented. Gemini had the greatest average concordance with multidisciplinary committee recommendations (3.89/5), with ChatGPT (3.33/5) and Grok (3.01/5) showing promise. Grok and Gemini concordance with multidisciplinary committee recommendations increased with positive nodal status when patients have limited options for management.
Conclusion:
Artificial intelligence large language models have substantial ability to replicate multidisciplinary committee recommendations but struggle with nuance. With improvement, artificial intelligence large language models can have a future role in health care decision support and guideline integration.
More Related Videos
05:33Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025
03:07Single-Port Robotic-assisted Transaxillary Breast-conserving Surgery: A Prospective, Single-arm, Non-randomized Phase IIa Clinical Trial
Published on: August 19, 2025