Related Experiment Video
Updated: Feb 28, 2026

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
Revisiting AI Interpretability in Precision Oncology: Why Predictive Accuracy Does Not Ensure Stable Feature
Souichi Oka1, Yoshiyasu Takefuji2
1Science Park Corporation, 3-24-9 Iriya-Nishi, Zama 252-0029, Japan.
Machine learning interpretability in oncology is often unreliable. This study introduces feature ranking consistency to ensure stable, trustworthy AI explanations for precision oncology, prioritizing stability alongside accuracy for clinical use.
Area of Science:
- Oncology
- Bioinformatics
- Computational Biology
Background:
- Artificial intelligence (AI) is increasingly vital in oncology for risk prediction, treatment planning, and biomarker discovery.
- Current AI evaluations often equate high predictive accuracy with reliable interpretation, potentially compromising reproducibility and clinical decision-making.
- This study addresses the need for robust interpretability metrics in AI for oncology.
Purpose of the Study:
- To reassess AI interpretability in oncology by introducing feature ranking order consistency as a stability-focused metric.
- To evaluate how AI model explanations respond to minimal input perturbations.
- To ensure AI models provide trustworthy and clinically actionable insights.
Main Methods:
- Compared supervised models (Linear Regression, LASSO, Random Forest, XGBoost) with unsupervised/statistical methods (PCA, Highly Variable Gene Selection, Spearman's rank correlation) using The Cancer Genome Atlas (TCGA) breast cancer multi-omics data.
- Assessed feature ranking stability by testing consistency after removing the top-ranked feature (<0.1% perturbation).
- Evaluated predictive performance using a Random Forest classifier with 10-fold cross-validation.
Main Results:
- Supervised models demonstrated unstable feature importance rankings even with minimal perturbations, indicating potentially fragile or misleading explanations despite high predictive accuracy.
- Unsupervised methods, specifically Highly Variable Gene Selection and Spearman's rank correlation, consistently produced stable and biologically coherent feature sets.
- These stable methods maintained competitive predictive performance compared to supervised approaches.
Conclusions:
- Interpretive instability is a significant limitation hindering the clinical application of many machine learning models in oncology.
- Integrating stability-based criteria, like feature ranking consistency, into AI evaluation frameworks is crucial for reproducible and trustworthy results.
- Prioritizing interpretability alongside accuracy is essential for the responsible and effective deployment of AI in precision oncology.
Related Concept Videos
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Receiver Operating Characteristic Plot
Improving Translational Accuracy
Improving Translational Accuracy
Accuracy and Precision
Biostatistics: Overview
Discrete variables are...

