Related Experiment Video
Updated: Jun 12, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Integrating large language models into prostate cancer training: evidence from comparative benchmarking and a pilot
Xuhao Liu1,2,3, Minfeng Chen4,5,6, Yi Cai7,8,9
1Department of Urology, Xiangya Hospital, Central South University, No.87 Xiangya Road, Changsha, Hunan Province, 410008, China.
BMC Medical Education
|June 11, 2026
Summary
Large language models (LLMs) improved urology resident knowledge in prostate cancer education, particularly for multiple-choice questions and research topics. Further studies are needed to confirm effectiveness in complex case analysis and generalizability.
Area of Science:
- Medical Education
- Artificial Intelligence in Medicine
- Urology Residency Training
Background:
- Limited evidence exists on the pedagogical quality and utility of large language models (LLMs) in urology residency training for prostate cancer education.
- Prostate cancer education within urology residency requires effective and modern teaching tools.
Purpose of the Study:
- To benchmark the performance of three LLMs in prostate cancer education.
- To evaluate the effectiveness of an AI-assisted teaching method in a pilot randomized trial for urology residents.
Main Methods:
- Phase 1: Benchmarked ChatGPT-4o, DeepSeek R1, and Gemini 2.0 using a 40-item prostate cancer question bank and expert ratings.
- Phase 2: Conducted a pilot randomized teaching trial with 34 urology residents, comparing an AI-assisted group with a control group receiving standard instruction.
- Utilized stratified block randomization and allocation concealment for the trial, with both groups receiving identical offline instruction.
Main Results:
- DeepSeek R1 demonstrated superior performance in expert ratings, especially for higher-order and innovation-oriented questions.
- The AI-assisted group achieved significantly higher closed-book examination scores compared to the control group (68.47 vs. 57.91, p=0.013).
- Improvements were most pronounced in Multiple-Choice Questions (MCQs) and research items, with no significant difference observed in Multidisciplinary Team (MDT) case analysis.
Conclusions:
- LLM-assisted teaching positively impacts knowledge-based examination performance in urology residents, particularly for MCQ-style and innovation-focused content.
- The effectiveness of LLMs in enhancing complex reasoning skills, such as MDT case analysis, remains uncertain.
- Carefully guided LLM integration may support residency education, but further multicenter studies are necessary to validate findings and ensure generalizability.