Related Experiment Video
Updated: Jul 15, 2026

07:41
Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
Published on: May 17, 2019
Auditing shortcut learning and misclassification in artificial intelligence-based breast cancer genomic subtyping
1Department of Computer Science, Boston University Metropolitan College, Boston, MA 02215, United States.
JAMIA Open
|April 6, 2026
Summary
This study developed a pseudo-Shapley additive explanations (SHAP) framework to audit AI models for breast cancer subtyping, revealing that models can over-rely on clinical predictors like HER2 status, mimicking shortcut learning and potentially impacting clinical decision support.
Area of Science:
- Computational biology and bioinformatics
- Artificial intelligence in healthcare
- Genomic data analysis
Background:
- AI models are increasingly used for breast cancer genomic subtyping.
- Concerns exist regarding AI models exhibiting 'shortcut learning,' over-relying on specific features.
- Developing interpretable frameworks to audit AI behavior is crucial for clinical trust.
Purpose of the Study:
- To simulate shortcut learning mechanisms in AI-based breast cancer genomic subtyping.
- To develop an interpretable and reproducible framework for auditing feature over-reliance.
- To implement this framework in a low-code environment using readily available clinical predictors.
Main Methods:
- Retrospective analysis of The Cancer Genome Atlas-Breast Invasive Carcinoma dataset (n=691).
- Trained a multinomial logistic regression model using clinical predictors (ER, PR, HER2, stage, age) to predict PAM50 subtypes.
- Implemented a pseudo-Shapley additive explanations (SHAP) method in Stata to estimate variable influence and validated against canonical SHAP values.
Main Results:
- The clinical-only model achieved moderate explanatory power (pseudo-R2=0.396).
- Progesterone receptor (PR) and HER2 status were significant predictors for Luminal B and HER2-enriched subtypes, respectively.
- Pseudo-SHAP identified disproportionate influence of PR and HER2 status (ΔP up to +0.29), with strong correlation to canonical SHAP values (Spearman r=0.91).
Conclusions:
- Shortcut learning was evident, with the AI model over-relying on surrogate clinical biomarkers (PR, HER2) instead of genomic signals.
- The pseudo-SHAP method provides a scalable and interpretable auditing framework for detecting shortcut learning in AI models.
- This approach can enhance transparency, reproducibility, and equity in AI for precision oncology, especially in resource-constrained settings.
Keywords:
artificial intelligencebiasbreast cancerexplainable AIgenomicsmachine learningmedical informatics
