Related Experiment Video
Updated: Jun 23, 2026

Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
Published on: May 17, 2019
ChatGPT Versus DeepSeek for Breast Cancer Information Retrieval: Quantitative Comparative Study.
Rima Hajjo1, Dima A Sabbah1, Sanaa K Bardaweel2
1Department of Pharmacy, Faculty of Pharmacy, Al-Zaytoonah University of Jordan, Airport Street, P.O. Box 130, Amman, 11733, Jordan, 962 64291511.
Artificial intelligence (AI) models like ChatGPT-4.0 and DeepSeek-V3 generate readable breast cancer information. DeepSeek-V3 showed higher accuracy and better citation alignment with experts, while ChatGPT-4.0 offered more consistent readability.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Oncology Content Generation
Background:
- Artificial intelligence (AI) is increasingly utilized for medical content creation.
- The clinical relevance and reliability of AI-generated medical information, particularly in complex fields like breast cancer, require further investigation.
- Evaluating AI performance in generating patient education materials is crucial for safe clinical integration.
Purpose of the Study:
- To compare the performance of ChatGPT-4.0 and DeepSeek-V3 in generating breast cancer information.
- To assess AI-generated content based on readability, clinical accuracy, completeness, clarity, depth, and citation reliability.
- To evaluate the alignment of AI-generated information with expert-validated answers and references.
Main Methods:
- Ten frequently asked questions on breast cancer were selected from patient education materials.
- ChatGPT-4.0 and DeepSeek-V3 generated 60 responses each for the selected questions.
- Expert reviewers rated responses on a 7-point Likert scale across five dimensions; readability was assessed using Flesch-Kincaid Grade Level; citation reliability was evaluated using Cohen κ and Fleiss κ statistics.
Main Results:
- AI-generated content was significantly more readable than expert references (mean Flesch-Kincaid Grade Level difference -2.60, P<.001).
- DeepSeek-V3 achieved higher content quality scores (mean 6.22) and demonstrated a statistically significant advantage in accuracy (P=.041) compared to ChatGPT-4.0 (mean 6.01).
- Both models exhibited high interrater agreement for citation reliability (Fleiss κ=0.842 for ChatGPT, 0.935 for DeepSeek), with DeepSeek-V3 showing stronger agreement with expert ratings (Cohen κ=0.782 vs 0.665 for ChatGPT).
Conclusions:
- Both ChatGPT-4.0 and DeepSeek-V3 produce readable and clinically relevant breast cancer information with comparable overall performance.
- DeepSeek-V3 demonstrated superior accuracy and better alignment with expert-validated citations, while ChatGPT-4.0 offered more consistent readability.
- Ongoing rigorous evaluation and quality assurance are essential for the responsible clinical application of AI-generated medical content.
More Related Videos
07:32Author Spotlight: Investigating Immune Cell Dynamics in the Tumor Microenvironment — Challenges and Innovations in Cancer Prognosis
Published on: April 12, 2024
07:13Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025