Related Experiment Video
Updated: Jun 16, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Benchmarking Large Language Models for Bladder Cancer: Comparative Knowledge Retrieval and Clinical Reasoning of
Lei Peng1,2,3, Yi Zhang1,2,3, Anguo Zhao3,4
1Department of Urology, The Second Affiliated Hospital of Kunming Medical University, Yunnan Institute of Urology, Kunming, People's Republic of China.
Annals of Surgical Oncology
|June 14, 2026
Summary
DeepSeek-R1 and ChatGPT-5 Pro show high accuracy in bladder cancer knowledge retrieval. While DeepSeek-R1 matches expert reasoning, it suggests unnecessary tests, unlike ChatGPT-5 Pro.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Knowledge Representation
- Clinical Decision Support Systems
Background:
- Evaluating the performance of advanced AI models in specialized medical domains is crucial.
- This study focuses on bladder cancer, a significant area within urology.
- Comparative analysis of AI models against human experts is essential for clinical integration.
Purpose of the Study:
- To compare the performance of DeepSeek-R1 and ChatGPT-5 Pro against human urologists.
- To assess AI capabilities in bladder cancer knowledge retrieval and clinical reasoning.
- To evaluate AI performance on standardized tests and complex clinical simulations.
Main Methods:
- A benchmark dataset of 91 multiple-choice questions and 3 real-world cases on bladder cancer was created.
- Five AI models, including DeepSeek (V3, R1) and OpenAI variants, were tested for accuracy and stability.
- DeepSeek-R1 and ChatGPT-5 Pro were benchmarked against senior urologists using a blinded expert panel evaluation.
Main Results:
- All AI models exceeded 92% accuracy on standardized tests, with ChatGPT-5 Pro achieving 97.52% and DeepSeek-R1 95.04%.
- In clinical simulations, DeepSeek-R1 showed comparable logical coherence to ChatGPT-5 Pro and human experts.
- Human urologists significantly outperformed AI in diagnostic test appropriateness, with DeepSeek-R1 recommending redundant tests.
Conclusions:
- DeepSeek-R1 exhibits strong accuracy and expert-level clinical reasoning, rivaling ChatGPT-5 Pro.
- While DeepSeek-R1 has minor limitations in readability and test selection, its reasoning is robust.
- AI models show promise in medical applications, but human oversight remains critical for optimal test appropriateness.
