Adversarial training and attribution methods enable evaluation of robustness and interpretability of deep learning

Flávio A O Santos1, Cleber Zanchettin1,2, Weihua Lei3

  • 1Centro de Informática, <a href="https://ror.org/047908t24">Universidade Federal de Pernambuco</a>, Recife, Pernambuco, 52061080, Brazil.

Physical Review. E
|December 18, 2024
PubMed
Summary

Adversarial training significantly alters deep learning model interpretability, making predictions more robust. This study benchmarks input attribution methods, revealing reliable approaches for trustworthy AI.