Related Experiment Video
Updated: Jun 16, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Benchmarking Large Language Models for Bladder Cancer: Comparative Knowledge Retrieval and Clinical Reasoning of
Lei Peng1,2,3, Yi Zhang1,2,3, Anguo Zhao3,4
1Department of Urology, The Second Affiliated Hospital of Kunming Medical University, Yunnan Institute of Urology, Kunming, People's Republic of China.
Background:
This study evaluates the comparative performance of DeepSeek-R1, ChatGPT-5 Pro, and urologists in the specific domains of bladder cancer knowledge retrieval and complex clinical reasoning.
Methods:
We constructed a benchmark dataset comprising 91 standardized multiple-choice questions on bladder cancer derived from MedQA, MedMCQA, and the Chinese National Medical Licensing Examination, alongside three retrospectively reconstructed real-world cases. Five advanced models, including DeepSeek (V3, R1) and OpenAI variants (ChatGPT-5, ChatGPT-5 Pro, ChatGPT-5 Mini), were evaluated. Accuracy and stability were assessed across three independent runs for standardized questions. In clinical simulations, DeepSeek-R1 and ChatGPT-5 Pro were benchmarked against human urologists. A blinded expert panel of three senior urologists evaluated responses using a 5-point Likert scale across four dimensions: readability, medical accuracy, diagnostic test appropriateness, and logical coherence.
Results:
In standardized testing, all models achieved >92% accuracy. ChatGPT-5 Pro ranked first (97.52%), followed closely by DeepSeek-R1 (95.04%), with both displaying superior stability. In clinical simulations, DeepSeek-R1 demonstrated logical coherence comparable to both human experts and ChatGPT-5 Pro (P > 0.05). However, DeepSeek-R1 scored significantly lower than ChatGPT-5 Pro regarding readability and "error-free" rates. Notably, human urologists significantly outperformed both AI models in diagnostic test appropriateness (P < 0.01), primarily because DeepSeek-R1 struggled with "test avoidance," tending to recommend redundant investigations (P < 0.0001).
Conclusions:
DeepSeek-R1 demonstrates excellent accuracy and expert-level clinical reasoning, exhibiting competitiveness with ChatGPT-5 Pro. Although slightly inferior in readability and prone to suggesting unnecessary tests, its core reasoning capabilities remain favorable.
