Related Experiment Video
Updated: Sep 9, 2025

09:47
Author Spotlight: Advancing Alzheimer's Research – Exploring Early Detection and Multi-Omics Approaches
Published on: December 15, 2023
1.2K
Improving the Interpretability of fMRI Decoding using Deep Neural Networks and Adversarial Robustness.
Patrick McClure1, Dustin Moraczewski2, Ka Chun Lam1
1Machine Learning Team, Functional Magnetic Resonance Imaging Facility, National Institute of Mental Health, Bethesda, MD, 20892, USA.
Aperture Neuro
|September 2, 2025
Summary
Deep neural networks (DNNs) are often uninterpretable. This study introduces new methods for creating reliable saliency maps to understand DNN predictions, particularly in neuroscience, showing adversarial training improves interpretability.
Area of Science:
- Neuroscience
- Machine Learning
- Data Science
Background:
- Deep neural networks (DNNs) are powerful tools for analyzing complex, high-dimensional data but are often considered "black boxes" due to their lack of interpretability.
- Understanding which input features drive DNN predictions is crucial for applications in cognitive neuroscience and neuroinformatics.
- Saliency maps are commonly used to visualize input feature importance, but existing methods can be unreliable, sensitive to noise, and difficult to evaluate.
Purpose of the Study:
- To review gradient-based saliency map methods for DNNs.
- To introduce a novel adversarial training method to enhance DNN robustness to input noise and improve interpretability.
- To propose and validate quantitative evaluation procedures for saliency map interpretability in neuroimaging data.
Main Methods:
- A review of existing gradient-based saliency map techniques.
- Development of an adversarial training approach for noise-robust DNNs.
- Introduction of two quantitative evaluation procedures for saliency maps, tested on synthetic data and functional magnetic resonance imaging (fMRI) data from the Human Connectome Project (HCP).
Main Results:
- Saliency maps generated by different methods exhibit significant variability in interpretability.
- Despite comparable decoding performance, DNN-derived saliency maps showed higher interpretability than those from linear models.
- The proposed adversarial training method resulted in saliency maps that outperformed alternative methods in interpretability.
Conclusions:
- The interpretability of saliency maps varies widely across different methods and model types.
- Adversarial training offers a promising approach to enhance the interpretability of DNNs in neuroimaging.
- The developed evaluation procedures provide a robust framework for assessing saliency map quality in decoding tasks.
