Related Experiment Video
Updated: Sep 9, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Safety and quality of public chatbots for lung cancer prognostic information: a comparative evaluation
Yanru Jiang1, Qianyun Wang1, Liang Zheng1
1Department of Thoracic Surgery, The First People's Hospital of Changzhou, Changzhou, Jiangsu, China.
Objective:
To compare the safety, accuracy, empathy, reliability, information quality, and readability of five publicly accessible large language model chatbots when answering patient-facing lung cancer prognostic questions under standardized single-turn English prompting.
Methods:
In this Chatbot Health Advice Reporting Transparency-guided cross-sectional evaluation, 53 standardized English prompts were submitted once to ChatGPT, Gemini, Copilot, DeepSeek, and Doubao through official web interfaces during April 1-21, 2026. Five blinded raters assessed 265 responses for safety, accuracy, empathy, DISCERN, EQIP, JAMA benchmark criteria, Global Quality Scale, and readability. Paired repeated-measures analyses were used.
Results:
Inter-rater agreement was good to excellent. Safety differed significantly across models (Cochran's Q = 14.089, df = 4, p = 0.007). Gemini generated the highest proportion of safe responses (48/53, 90.6%), whereas DeepSeek generated the lowest (33/53, 62.3%). The only adjusted pairwise safety difference that remained significant was Gemini versus DeepSeek (adjusted p = 0.023). Accuracy, empathy, reliability, information quality, and readability also differed significantly across models (all p < 0.001). Gemini showed the most favorable descriptive profile for safety, accuracy, empathy, and reliability, while Copilot produced the most readable responses.
Conclusion:
Public-facing chatbots differed substantially in safety, reliability, communication quality, and readability. These findings are time-, interface-, and prompt-dependent. Chatbots may support general patient education but should not replace individualized clinician-led prognostic communication.