Related Experiment Video
Updated: Sep 19, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance evaluation of large language models in bladder cancer patient education Q&A: a cross-sectional study
Dian Wan1, Youwen Li1, Zheng Dong1
1Department of Urology, Huanggang Central Hospital, Affiliated Huanggang Hospital, Hubei University of Science and Technology, Huanggang, China.
Background:
Bladder cancer ranks among the most prevalent urological tumors worldwide, with its global incidence continuing to rise steadily. Although patient education materials (PEMs) play a crucial role in enhancing disease comprehension and supporting joint clinical decision-making, current online resources frequently surpass the readability thresholds recommended for the general public. Large language models (LLMs) hold promise for health communication, yet no systematic assessment has been conducted regarding their feasibility and trustworthiness specifically for bladder cancer patient education.
Objective:
This study aimed to systematically benchmark five leading LLMs in producing question-and-answer content for bladder cancer science popularization, with a particular focus on readability, informational quality, and appropriateness for patient education.
Methods:
In this cross-sectional simulation study, 20 common patient questions covering five disease domains were compiled. On January 15, 2026, each question was submitted identically to five publicly available LLMs (Doubao, DeepSeek, Kimi, Gemini, and ChatGPT). Readability was evaluated using seven conventional metrics. Two independent pharmacists, blinded to model identity, rated the responses using the Chinese version of the Patient Education Materials Assessment Tool for print materials (C-PEMAT-P) and the Global Quality Score (GQS). Additionally, two independent clinical specialists assessed factual accuracy and alignment with the Chinese Bladder Cancer Diagnosis and Treatment Guidelines (2024 edition) employing a 4-point scale. Cohen's kappa was used to determine inter-rater reliability.
Results:
ChatGPT, DeepSeek, and Doubao outperformed Kimi and Gemini on both C-PEMAT and GQS (all P < 0.001), indicating superior understandability, actionability, and overall quality. Across all models, median C-PEMAT scores ranged from 8 to 10, suggesting broadly acceptable suitability for patient education. Readability varied significantly by content domain, with treatment-oriented texts showing the highest complexity. ChatGPT achieved the best alignment with clinical guidelines. No model produced harmful advice or directly contradicted guideline recommendations. Traditional readability measures correlated weakly with GQS, whereas C-PEMAT showed a moderate positive correlation (r = 0.34).
Conclusion:
Current mainstream LLMs demonstrate initial potential for generating educational content on bladder cancer, albeit with considerable heterogeneity across models. Disease-specific evaluation instruments for patient education materials are more effective than general readability formulas in reflecting perceived quality. Our results advocate for a prudent, assistive role of LLMs in health communication under a human-AI collaborative model.