Related Experiment Video
Updated: Jan 10, 2026

Studying Triple Negative Breast Cancer Using Orthotopic Breast Cancer Model
Published on: March 20, 2020
The Accuracy And Clinical Relevance of Chat GPT-4 in Triple Negative Breast Cancer Research
Ramakrishna Gummadi1, Sai Kiran S S Pindiprolu1, Chirravuri S Phani Kumar1
1Aditya Pharmacy College, Surampalem, Andhrapradesh, 533437, India.
Background:
Triple negative breast cancer (TNBC) is an aggressive subtype of breast cancer characterized by the lack of estrogen receptor(ER), progesterone receptor (PR), and human epidermal growth factor receptor 2 (HER2). The absence of these receptors reduces the effectiveness of targeted treatment approaches. With the increasing use of artificial intelligence (AI) in medical research and clinical decision- making, there is growing interest in evaluating the accuracy and reliability of large language models (LLMS), such as chatGPT-4, in oncology related applications.
Objective:
The research aims to systematically assess the reliability of ChatGPT-4 in addressing frequently asked questions related to TNBC in four critical areas: diagnosis, treatment, prognosis and survival, and quality of life. Expert evaluations and statistical analyses are employed to measure the accuracy of the models responses.
Methods:
A set of 100 questions related to TNBC was gathered from credible medical sources, including peer-reviewed journals and clinical oncology specialists evaluated the response generated by ChatGPT-4 using a structured assessment framework, classifying each answer into one of four accuracy levels, completely inaccurate, partially accurate, accurate but lacking depth and highly accurate.
Results:
To evaluate the consistency among reviewers, Cohen's kappa coefficient was calculated, and descriptive statistical analysis was conducted to identify overall accuracy patterns. The findings indicated that 73% of the responses were classified as either "Accurate" or "Highly Accurate", suggesting the potential of ChatGPT-4 as a supplementary resource for obtaining information on TNBC. However, 27% of the responses were categorized as "partially accurate" or 'Completely Inaccurate, "highlighting gaps in contextual understanding and instances of misinformation.. Cohen's kappa coefficient was recorded at 0.007, reflecting a week level of agreement among evaluators and highlighting the impact of subjective interpretation. The model demonstrated strong performance in well-established areas such as chemotherapy protocols and diagnostic procedures but faced challenges with emerging research topics, personalized treatment recommendations, and fertility related concerns.
Conclusion:
ChatGPT-4 exhibits significant potential in summarizing information on TNBC; however, the accuracy of its responses varies depending on the complexity and specificity of the queries. Due to inconsistencies and low inter-rater reliability, AI-generated medical content requires verification by medical professionals before being applied to patient care or clinical decision-making. Future developments in large language models should focus on reducing inaccuracies, incorporating the latest medical data, and improving adaptability to better support personalized medicine.
Insights
ChatGPT-4 shows promise for triple-negative breast cancer (TNBC) information, but accuracy varies. Medical professionals must verify AI-generated content due to potential inaccuracies and low reliability.
Area of Science:
- Oncology
- Artificial Intelligence in Medicine
- Medical Informatics
Background:
- Triple-negative breast cancer (TNBC) is an aggressive subtype lacking ER, PR, and HER2 receptors, limiting targeted therapies.
- Artificial intelligence (AI) and large language models (LLMs) like ChatGPT-4 are increasingly explored for oncology applications.
- Evaluating LLM reliability in medical contexts is crucial for safe and effective implementation.
Purpose of the Study:
- To systematically assess ChatGPT-4's reliability for frequently asked questions on TNBC.
- To evaluate accuracy across diagnosis, treatment, prognosis, and quality of life domains.
- To measure the accuracy of AI-generated responses through expert evaluation and statistical analysis.
Main Methods:
- A curated set of 100 TNBC-related questions from credible medical sources was used.
- ChatGPT-4 responses were evaluated by clinical oncology specialists using a structured framework.
- Responses were classified into four accuracy levels: completely inaccurate, partially accurate, accurate but lacking depth, and highly accurate.
Main Results:
- 73% of ChatGPT-4 responses were rated as "Accurate" or "Highly Accurate".
- 27% of responses were "partially accurate" or "Completely Inaccurate," indicating potential misinformation.
- Low inter-rater reliability (Cohen's kappa = 0.007) suggests subjective interpretation challenges.
Conclusions:
- ChatGPT-4 has potential as a supplementary TNBC information resource, but accuracy is variable.
- AI-generated medical content requires rigorous verification by healthcare professionals.
- Future LLMs need improved accuracy, updated data integration, and better adaptability for personalized medicine.

