Related Experiment Videos
Large Language Models for Breast Cancer Education: A Comparative Analysis of Quality, Reliability and Readability
1Department of General Surgery, Istanbul Research and Training Hospital, Istanbul, Turkey.
Abstract:
BackgroundPatients increasingly consult artificial intelligence (AI) tools for breast cancer information. While Large Language Models (LLMs) enhance information accessibility, their accuracy, reliability, and alignment with patient health literacy remain critical concerns. This study compared the quality, reliability, and readability of breast cancer-related responses generated by ChatGPT and Gemini.MethodsIn this cross-sectional study, conducted between March 20 and March 31, 2026, 40 questions spanning diagnosis, treatment, genetics, and follow-up were submitted to ChatGPT-5.3 and Gemini 3.0 Flash. Three independent surgeons evaluated the responses in a double-blinded manner using the modified DISCERN (mDISCERN) for reliability and Global Quality Score (GQS) for content quality assessment. Readability was assessed via Flesch Reading Ease (FRES), Flesch-Kincaid Grade Level (FKGL), Gunning Fog Index (GFI), and Simple Measure of Gobbledygook (SMOG) indices.ResultsGemini demonstrated statistically significant superiority over ChatGPT in both GQS (4.35 ± 0.30 vs 3.90 ± 0.29; P < .001) and mDISCERN (3.63 ± 0.58 vs 3.02 ± 0.48; P < .001) scores. In readability analysis, Gemini exhibited higher FRES (53.90 vs 45.13) and lower FKGL (9.31 vs 10.91) values, indicating enhanced patient accessibility (P < .05). For both models, the "Diagnosis" category yielded the highest readability, whereas "Treatment" scored the lowest. Inter-rater reliability for mDISCERN was moderate (ICC = 0.595).ConclusionsGemini significantly outperforms ChatGPT in response quality, reliability, and linguistic accessibility for breast cancer education. However, both models exceed the recommended sixth-grade reading level, indicating suboptimal optimization for general health literacy. While LLMs serve as promising auxiliary tools, expert supervision and cross-validation remain mandatory to ensure patient safety.