Related Experiment Video
Updated: Jun 11, 2026

06:32
Multiplex Immunohistochemical Analysis of the Spatial Immune Cell Landscape of the Tumor Microenvironment
Published on: August 18, 2023
Beyond attention heatmaps: How to get better explanations for multiple instance learning models in histopathology.
Mina Jamshidi Idaji1, Julius Hense1, Tom Neuhäuser1
1Berlin Institute for the Foundations of Learning and Data, Berlin, Germany; Machine Learning Group, Technische Universität Berlin, Berlin, Germany.
Medical Image Analysis
|June 9, 2026
Summary
We developed a framework to evaluate multiple instance learning (MIL) heatmaps in computational histopathology. Perturbation, LRP, and IG methods show superior performance for validating MIL models and discovering biomarkers.
Area of Science:
- Computational pathology
- Artificial intelligence in medicine
- Digital pathology
Background:
- Multiple instance learning (MIL) is crucial for analyzing gigapixel whole slide images in computational histopathology.
- Heatmaps are commonly used for MIL model validation and biomarker discovery, but their reliability is under-explored.
- Evaluating heatmap quality is essential for trustworthy AI in pathology.
Purpose of the Study:
- Introduce a general framework to assess the quality of MIL heatmaps without additional labels.
- Benchmark six explanation methods across diverse histopathology tasks, MIL architectures, and patch encoders.
- Demonstrate the utility of high-quality MIL heatmaps for biological validation and insight discovery.
Main Methods:
- Developed a novel framework for evaluating MIL heatmap quality.
- Conducted a large-scale benchmark experiment comparing six explanation methods.
- Assessed methods across classification, regression, and survival tasks, and various MIL/encoder architectures.
Main Results:
- Explanation quality is significantly influenced by MIL model architecture and task type.
- Perturbation ('Single'), Layer-wise Relevance Propagation (LRP), and Integrated Gradients (IG) methods outperformed others.
- Attention-based and gradient-based saliency heatmaps often failed to accurately reflect model decision-making.
- Demonstrated correlation of MIL heatmaps with spatial transcriptomics and discovered distinct HPV prediction strategies.
Conclusions:
- Validated MIL heatmaps are critical for reliable model assessment and biological discovery in digital pathology.
- The proposed framework and best-performing methods (Perturbation, LRP, IG) enhance explainability and trustworthiness of MIL models.
- Promotes broader adoption of explainable AI (XAI) for advancing digital pathology research and clinical applications.
