Related Experiment Video
Updated: Aug 20, 2026

A New Technique for Treating Low-risk Prostate Cancer—Super Active Surveillance
Published on: November 7, 2025
ChatGPT-4o as a decision-support tool in a urological tumour board: a prospective evaluation
Javier De la Torre-Trillo1, Alberto Zambudio Munuera2, Mayte Delgado Ureña1,3
1Urology Department, Hospital Regional Santa Ana de Motril, Av. Enrique Martín Cuevas S/N, 18600, Motril, Granada, Spain.
Purpose:
To evaluate ChatGPT-4o in a real-world urological multidisciplinary tumour board (MTB), with concordance for the final clinical recommendation as the protocol-defined primary comparison.
Methods/Patients:
Prospective study of 106 consecutive MTB discussions in 92 patients (14 discussed twice) at Hospital Regional Santa Ana de Motril, Spain (January-December 2025). ChatGPT-4o produced a guideline-informed decision with rationale from de-identified summaries, revealed only after MTB consensus. Seven members rated outputs on a five-level correctness scale (operationally defined high concordance, scores ≥4). Benchmarks were ≥80% recommendation concordance, ≥10% added value, ≥5% attitude change. Contemporaneous evaluator notes were reviewed post hoc for documented model errors.
Results:
High concordance was 91.5% (summarisation), 84.0% (reasoning) and 79.2% (recommendation; 95% CI, 70.6-85.9), below benchmark (Friedman χ2 = 34.756; p < 0.001; κ = 0.618). Results were consistent at patient level (recommendation 77.2%; κ = 0.585). Added value was 13.2% (higher in metastatic disease, p = 0.011); no attitude change occurred (0/106). The post hoc review documented model errors in 19/106 discussions (17.9%), including 13/84 (15.5%) discussions rated highly concordant and 6/22 (27.3%) rated below that threshold; four errors (3.8%) were graded as potentially harmful.
Conclusions:
Final recommendation concordance did not meet the protocol-defined benchmark, ChatGPT-4o never altered an MTB decision, and clinically relevant errors occurred even among highly concordant outputs, showing that concordance alone does not guarantee safety. These findings suggest potential utility for case summarisation and guideline-oriented reasoning requiring independent blinded validation, rather than a defined clinical role at this stage.
