Related Experiment Videos
Explainable deep learning improves human mental models of self-driving cars
Eoin M Kenny1, Akshay Dharmavaram2, Sang Uk Lee2
1Computer Science and Artificial Intelligence Laboratory (CSAIL), Massachusetts Institute of Technology, Cambridge, MA, USA. ekenny@mit.edu.
Abstract:
Self-driving cars increasingly rely on deep neural networks to achieve human-like driving1-3. The opacity of these black-box planners makes it challenging to accurately anticipate when they will fail4-6, with potentially catastrophic consequences7-9. Although research into interpreting these systems has surged, most of it is confined to simulations or toy setups because of the difficulty of real-world deployment10,11, leaving the practical utility of these techniques unknown. Here, we introduce the Concept-Wrapper Network (CW-Net), a method for faithfully explaining the behaviour of machine-learning-based planners that causally grounds their reasoning in human-interpretable concepts without sacrificing performance. We deploy CW-Net on a real self-driving car and show that the resulting explanations improve the human driver's mental model of the vehicle, allowing them to better predict its behaviour, particularly in surprising situations. This demonstrates that explainable deep learning integrated into self-driving cars can be both understandable and useful in a realistic deployment setting. We anticipate our method could be applied to other safety-critical systems, such as autonomous drones and robotic surgeons, as well as to other architectures, such as end-to-end learning systems and vision-language-action models. Overall, our study establishes a deployment-validated pathway to interpretability for autonomous agents, which could help make them more transparent and safe.