Explainability Methods for AI-Assisted Diagnosis of Lymph Node Metastases in Digital Pathology: A Quantitative
1Department of Electrical Engineering, Pontifical Catholic University of Rio de Janeiro (PUC-Rio), Rio de Janeiro 22451-900, Brazil.
Diagnostics (Basel, Switzerland)
|June 26, 2026
Summary
Explainable AI (XAI) methods for detecting cancer in pathology images were evaluated. GradCAM++ showed best spatial agreement, SHAP offered highest faithfulness, and squaregrid LIME provided a cost-effective baseline for clinical use.
Area of Science:
- Digital pathology
- Artificial intelligence in medicine
- Explainable AI (XAI)
Background:
- AI systems for detecting lymph node metastases in histopathological images show near-expert performance but lack clinical transparency.
- This opacity hinders clinical adoption and regulatory approval of AI tools in pathology.
- A quantitative framework is needed to evaluate and compare XAI methods for digital pathology.
Purpose of the Study:
- To present the first rigorous quantitative framework for evaluating and comparing XAI methods in digital pathology.
- To provide evidence-based guidance for the clinical deployment of explainable AI in histopathology.
- To assess the spatial agreement and faithfulness of different XAI techniques.
Main Methods:
- Four XAI techniques (LIME, GradCAM, GradCAM++, SHAP) were applied to three convolutional neural networks (VGG19, ResNet50, EfficientNetB3) trained on the PatchCamelyon (PCam) benchmark.
- Quantitative evaluation used spatial agreement metrics (IoU, Sørensen-Dice) with pathologist annotations and faithfulness metrics (AOPC, insertion/deletion AUC).
- Threshold sensitivity analysis was performed across different binarization thresholds, including Otsu automatic thresholding.
Main Results:
- GradCAM++ achieved the highest spatial agreement with pathologist annotations (mean IoU = 0.52 ± 0.14).
- SHAP (via DeepExplainer) yielded the highest faithfulness scores (AOPC = 0.61 ± 0.08).
- The parameter-free squaregrid LIME variant offered a favorable trade-off at significantly lower computational cost compared to LIME AVG.
Conclusions:
- GradCAM++ is recommended for high-throughput clinical workflows due to its spatial accuracy.
- SHAP is suitable for research requiring maximal faithfulness in AI explanations.
- Squaregrid LIME serves as a transparent, parameter-free baseline for clinical communication and audits, outperforming LIME AVG.


