Related Experiment Video
Updated: May 5, 2026

Author Spotlight: Advancing Early Detection and Treatment of Gastrointestinal Tumors
Published on: February 16, 2024
Assessing the Accuracy and Readability of Generative Artificial Intelligence Responses for Esophageal and Gastric
Shanhu Ran1,2, Wenlong Guan2,3, Ran Wei2,4
1Department of Intensive Care Unit, Sun Yat-sen University Cancer Center, Guangzhou 510060, China.
Abstract:
Background: Generative artificial intelligence (GenAI) models are increasingly used for medical information retrieval, due to their accessibility and efficiency. However, the accuracy and readability of their responses, specifically for upper gastrointestinal cancers, remain inadequately evaluated. This gap highlights the need for rigorous assessment to ensure reliable patient education and clinical integration. Objective: This study aimed to assess the accuracy and readability of responses generated by four prominent GenAI models (Kimi, DeepSeek, ChatGPT, and Gemini) when addressing patient-focused questions related to esophageal and gastric cancers. Methods: Twenty-five standardized medical questions about esophageal and gastric cancer covering domains of disease definition, treatment and management were posed to each model. Responses were assessed by four oncologists for accuracy by a 5-point Likert scale and analyzed for readability using Flesch-Kincaid Reading Ease, Flesch-Kincaid Grade Level, and SMOG metrics. High-interest questions for patients were identified via questionnaires. Results: Comparing the accuracy of GenAI-generated responses, DeepSeek achieved the highest overall accuracy score and outperformed other models in questions about definitions and treatments, while ChatGPT excelled in management-related inquiries. In subgroup analysis, GenAI models exhibited higher accuracy in answering definition and management questions, which patients preferred to inquire, compared with questions about cancer therapies. The responses produced by all models required a reading capacity from 11th-grade to college level. Conclusions: This study revealed that in this comparative evaluation application of GenAI models, DeepSeek provides the most accurate responses for upper GI cancer inquiries about definition and treatment, while ChatGPT showed superiority in management-related questions. However, all models generate texts requiring advanced reading levels, highlighting a need for readability optimization without compromising accuracy. GenAI shows promise for patient education but requires rigorous validation for clinical integration.
Related Concept Videos
Assessment of the Gastrointestinal System I: Subjective Data
Health History
The initial step in assessing the GI system is obtaining a comprehensive health history. This includes inquiring about the patient's history or presence of problems...
Esophageal Achalasia

