Related Experiment Videos
Artificial Intelligence Chatbots as Sources of Cancer Pain Information: A Comparative Evaluation of Quality,
Qianpeng Li1, Shuai Zhang2, Xiao Ma3
1Department of Hematology, Weifang People's Hospital, Weifang, 261041, People's Republic of China.
Purpose:
To compare the quality, transparency, educational value, and readability of cancer pain information generated by four AI chatbots and to assess whether the outputs met prespecified readability benchmarks for patient education.
Methods:
This online cross-sectional comparative study was conducted on July 22, 2026. Nine unmodified Google-Trends-derived cancer-pain-related queries were submitted to ChatGPT 5.5, Microsoft Copilot, Google Gemini 3.5 Flash, and Perplexity, yielding 36 responses. Two oncology clinicians independently assessed the responses using DISCERN, Ensuring Quality Information for Patients (EQIP), Journal of the American Medical Association (JAMA) benchmark criteria, and the Global Quality Score (GQS). Six established indices assessed readability. Matched model comparisons used Friedman tests with query as the repeated unit and Kendall's W as the omnibus effect size; significant quality outcomes were followed by Holm-adjusted paired Wilcoxon signed-rank tests.
Results:
DISCERN did not differ significantly across models (χ2(3)=6.682, P=0.083, W=0.247). EQIP differed across models (χ2(3)=20.721, P<0.001, W=0.767), with Copilot scoring higher than ChatGPT, Gemini, and Perplexity (Holm-adjusted P=0.023 for each). JAMA also differed (χ2(3)=19.645, P<0.001, W=0.728); Copilot and Perplexity scored higher than Gemini (Holm-adjusted P=0.023 and.039, respectively). GQS differed overall (χ2(3)=10.500, P=0.015, W=0.389), although no pairwise comparison remained significant after Holm adjustment. All six readability indices differed overall across models after Holm correction; descriptively, Gemini showed the greatest estimated reading difficulty. Model medians exceeded the prespecified grade-level benchmarks, and FRES medians were below 80.
Conclusion:
The four chatbots differed across information quality, source transparency, educational utility, and formula-based readability. Claim-level clinical accuracy and safety were not evaluated in this study and warrant separate guideline-based assessment.
Related Concept Videos
Cancer Survival Analysis
Mouse Models of Cancer Study
The development of transgenic, knockout, and knock-in mice has led to an exponential increase in their use as model organisms in research,...
Combination Therapies and Personalized Medicine
The combination of the drug acetazolamide and sulforaphane is a good example of combination therapy to treat cancer. The cells in the interior of a large tumor often die due to the hypoxic and...