Related Experiment Video
Updated: Sep 30, 2026

Patient-derived Orthotopic Xenograft Models for Human Urothelial Cell Carcinoma and Colorectal Cancer Tumor Growth and Spontaneous Metastasis
Published on: May 12, 2019
Development of a ChatGPT-based assistant in multidisciplinary team discussions in oncological urology: the UroMDT
Riccardo Lombardo1, Maria Giovanna Stagliano2, Marta Santioni2
1Department of Urology, Ospedale Sant'Andrea, Sapienza, University of Rome, Sapienza Università Di Roma, Viale Di Grottarossa, 1089 00189, Rome, Italy. riccardo.lombardo@uniroma1.it.
Introduction:
This study aimed to develop and evaluate a ChatGPT-based application designed to assist multidisciplinary team (MDT) discussions in oncological urology.
Patients And Methods:
A ChatGPT-based tool (UroMDT Advisor) was developed to: (1) generate case summaries, (2) provide evidence-based treatment suggestions aligned with EAU, ESMO, and NCCN guidelines, (3) offer supporting references, (4) compare opinions across specialties to identify inconsistencies, (5) draft final MDT reports, and (6) create patient summaries. A consecutive series of anonymized oncological cases from eight centers were entered into the system. For each case, GPT-generated opinions were compared with those of MDT participants (urologists, oncologists, radiotherapists) and the final consensus decision. The model operated in a blinded fashion during concordance testing and was subsequently unblinded to produce final reports and patient summaries. Two expert urologists independently assessed the accuracy, completeness, and clarity of GPT outputs.
Results:
A total of 180 cases were analyzed: 123 (68.3%) prostate, 21 (11.7%) renal, 21 (11.7%) urothelial, 12 (6.7%) testicular, and 3 (1.7%) penile cancer cases. Exact agreement between UroMDT Advisor and urologists, oncologists, radiation oncologists, and final MDT decisions was 61.7%, 66.7%, 63.3%, and 66.7%, respectively; the corresponding unweighted Cohen's kappa values were 0.401, 0.489, 0.441, and 0.482. After resolution of discordant ratings, the GPT-generated final reports were rated as accurate in 90% of cases, fairly accurate in 8%, and inaccurate in 2%. Completeness was rated high in 95%, moderate in 2%, and low in 3% of cases, while clarity was rated high in 95% and moderate in 5%.
Conclusions:
The UroMDT Advisor showed moderate agreement with final MDT decisions and produced highly rated outputs. These findings support further evaluation of large language models as supervised decision-support and documentation tools in oncological MDT workflows while maintaining human oversight.