Related Experiment Video
Updated: Mar 8, 2026

Optimization of a Multiplex RNA-based Expression Assay Using Breast Cancer Archival Material
Published on: August 1, 2018
Analysis of Large Language Model Decision Making in Hormone Receptor-Positive/Human Epidermal Growth Factor Receptor
Roberto Buonaiuto1,2,3, Aldo Caltavituro2,3, Rossana Di Rienzo1,2
1Department of Breast and Thoracic Oncology, Istituto Nazionale Tumori IRCCS "Fondazione G. Pascale," Naples, Italy.
GPT-4o showed modest agreement with clinicians on adjuvant treatment for early breast cancer before genomic testing, but improved significantly after Oncotype DX results were available. This highlights the potential of large language models as clinical decision-support tools.
Area of Science:
- Oncology
- Artificial Intelligence
- Genomics
Background:
- Hormone receptor-positive (HR+)/human epidermal growth factor receptor 2-negative (HER2-) early breast cancer treatment decisions often rely on multigene assays like Oncotype DX.
- Large language models (LLMs) are emerging as potential tools to aid clinical decision-making.
Purpose of the Study:
- To evaluate GPT-4o's accuracy in adjuvant treatment recommendations for HR+/HER2- early breast cancer.
- To compare GPT-4o's recommendations with clinician decisions, utilizing Oncotype DX data.
- To explore GPT-4o's utility as a clinical decision-support tool.
Main Methods:
- A comparative analysis of clinician and GPT-4o treatment recommendations (chemotherapy + endocrine therapy vs. endocrine therapy alone) was performed.
- Data from 607 patients (cohort 1) and 237 patients (cohort 2) were analyzed, including pre- and post-Oncotype DX results.
- Agreement was assessed using rates and Cohen's kappa; Oncotype DX accuracy was evaluated using AUC.
Main Results:
- Pre-test agreement between clinicians and GPT-4o was modest (68% in cohort 1, 70% in cohort 2).
- Clinicians were more likely to recommend chemotherapy before Oncotype DX results compared to GPT-4o.
- Post-test agreement significantly improved to 93% (cohort 1) and 90% (cohort 2), with GPT-4o demonstrating higher accuracy in predicting genomic risk categories.
Conclusions:
- GPT-4o's adjuvant treatment recommendations showed modest pre-test concordance with clinicians but substantial post-test agreement.
- Multigene testing (Oncotype DX) is crucial for refining treatment decisions.
- LLMs like GPT-4o show promise as valuable decision-support tools in routine oncology practice.
More Related Videos
08:28Validated Immunochemical Assay for Comprehensive Determination of the Human Epidermal Growth Factor Receptor 2 Released from and Bound to Cells
Published on: May 9, 2025
07:41Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
Published on: May 17, 2019