Benchmarking Large Language Models Against Multidisciplinary Tumor Boards in Urological Oncology: Results from the
Emily Rinderknecht1, Maximilian Haas1, Marco J Schnabel1
1Department of Urology, St. Josef Medical Center, University of Regensburg, Regensburg, Germany.
Background And Objective:
Multidisciplinary tumor boards (MTBs) are the gold standard for oncological treatment planning, but their implementation is resource intensive. Large language models (LLMs) such as ChatGPT-4 and Claude 3.5 Sonnet have emerged as scalable tools for clinical decision support. As their comparative performance in urological oncology remains largely untested in prospective trials, this study aimed to prospectively evaluate them against real MTBs.
Methods:
In this prospective, multicenter, noninferiority study (DRKS00034797), we evaluated whether therapeutic recommendations by ChatGPT-4 and Claude 3.5 Sonnet were noninferior to those of MTBs across 110 representative case scenarios involving locally advanced or metastatic genitourinary cancer. Standardized prompts elicited recommendations from both LLMs, which were rated independently by two blinded uro-oncologists using the validated modified System Causability Scale (mSCS). The predefined noninferiority margin was 0.15 mSCS points.
Key Findings And Limitations:
The mean mSCS score of the MTBs was 0.849 (standard deviation [SD] = 0.157), setting the noninferiority margin at 0.699. Claude 3.5 Sonnet scored 0.731 (SD = 0.178; 95% confidence interval [CI]: 0.697-0.765), narrowly missing noninferiority. ChatGPT-4 scored 0.660 (SD = 0.193; 95% CI: 0.623-0.696), clearly below the margin. A subgroup analysis revealed better LLM performance in locally advanced versus metastatic cases (p < 0.05). Limitations include the use of synthetic cases and inherent LLM output variability.
Conclusions And Clinical Implications:
Neither LLM matched the MTB standards. These findings highlight the current limitations of public LLMs, or of their use in our study, in supporting complex oncological decisions. This underscores the need for further validation and contextual refinement before integration into multidisciplinary care.
More Related Videos
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
09:53Quantifying the Brain Metastatic Tumor Micro-Environment using an Organ-On-A Chip 3D Model, Machine Learning, and Confocal Tomography
Published on: August 16, 2020
