Related Experiment Video
Updated: May 13, 2025

08:51
Author Spotlight: Integrated Multi-Omics Analysis for Unveiling Multicellular Immune Signatures in Clinical Heart Attack Cohorts
Published on: September 20, 2024
1.0K
BenchXAI: Comprehensive benchmarking of post-hoc explainable AI methods on multi-modal biomedical data.
Jacqueline Michelle Metsch1, Anne-Christin Hauschild2
1Institute for Medical Informatics, University Medical Center Göttingen, Germany.
Computers in Biology and Medicine
|April 16, 2025
Summary
Benchmarking explainable artificial intelligence (AI) methods is crucial for medical applications. BenchXAI evaluates fifteen AI explainability techniques across diverse biomedical data, identifying top-performing methods for reliable medical AI.
Area of Science:
- Biomedical data analysis
- Artificial intelligence in medicine
- Explainable AI (XAI)
Background:
- Digitalization and AI offer predictive modeling opportunities in medicine.
- Deep learning models excel but suffer from a 'black-box' problem, hindering trust.
- Existing explainable AI (XAI) methods are often evaluated on single data types, limiting cross-modal applicability.
Purpose of the Study:
- To develop and introduce BenchXAI, a novel benchmarking package for evaluating XAI methods in biomedical contexts.
- To assess the robustness, suitability, and limitations of fifteen XAI methods across diverse biomedical data modalities.
- To provide a framework for the statistical evaluation and visualization of XAI method performance and robustness.
Main Methods:
- Development of BenchXAI, a package for comprehensive XAI method evaluation.
- Application of BenchXAI to three common biomedical tasks: clinical, imaging/signal, and biomolecular data.
- Implementation of a novel sample-wise normalization for post-hoc XAI methods to enable statistical analysis.
Main Results:
- Four XAI methods (Integrated Gradients, DeepLift, DeepLiftShap, GradientShap) demonstrated strong performance across all evaluated biomedical tasks.
- Certain XAI methods (Deconvolution, Guided Backpropagation, LRP-α1-β0) showed limitations on specific tasks.
- Benchmarking revealed performance variations among XAI methods depending on the data modality and task.
Conclusions:
- BenchXAI provides a vital resource for comparing and validating XAI methods in medicine.
- The study highlights the need for modality-aware XAI evaluation for reliable medical AI.
- Findings support the increasing necessity of explainable AI in the biomedical domain, especially with emerging regulations like the EU AI Act.

