Related Experiment Video
Updated: Aug 21, 2026

Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025
Multi-view explainability and ensemble deep learning for prostate lesion classification: A comprehensive study
Mehmet A Gulum1, Mehmed Kantardzic2, Christopher M Trombley3
1Department of Computer Engineering, Hitit University, Çorum, 19030, Turkiye.
Background And Objective:
Explainable artificial intelligence is essential for clinical adoption of deep learning models in prostate magnetic resonance imaging. Although ensemble learning can improve robustness, its impact on explanation stability, spatial consistency, and clinical interpretability remains insufficiently quantified. This study aims to evaluate post-hoc interpretability methods across both single-model and ensemble configurations, and to examine whether ensemble-based explanations provide more reliable and clinically meaningful insights than single-model explanations. Critically, this work treats interpretability as a measurable property rather than a purely qualitative visualization.
Methods:
Convolutional neural networks and bagging-based ensemble models (five VGG16-based classifiers trained on bootstrap samples with replacement, aggregated by soft averaging) were trained on the public PROSTATEx dataset using T2-weighted and apparent diffusion coefficient images. Visual explanations were generated using Gradient-weighted Class Activation Mapping (Grad-CAM) and saliency maps. Lesion localization was evaluated using centroid distance and Dice similarity coefficient with expert-annotated lesion masks. An agreement metric was introduced to quantify spatial consistency between attribution methods and its relationship with prediction reliability.
Results:
The baseline classifier achieved an area under the curve of 0.84, with sensitivity of 0.81 and specificity of 0.86. Grad-CAM localized lesion centroids with higher precision on T2-weighted images (mean error 6.93 pixels) than apparent diffusion coefficient images (mean error 16.3 pixels). Combining saliency maps and Grad-CAM improved the mean Dice score from 0.42/0.45 (individual methods) to 0.52. Ensemble-based explanations were significantly smoother and less variable than individual classifier explanations (Mann-Whitney U, Levene test, all p<0.001). The agreement metric strongly separated correctly and incorrectly classified cases (Mann-Whitney U=84.5, p<0.001; point-biserial r=0.749; ROC-AUC =0.960).
Conclusions:
These findings suggest that interpretability quality can be quantitatively assessed and improved through multi-method and ensemble-based analysis. The proposed agreement-driven framework enhances explanation robustness and supports reliable and transparent clinical decision support for prostate magnetic resonance imaging.