Related Experiment Videos
Machine learning and SHAP-based explainability for accident severity prediction: evidence from a strategic road
Rodrigo Aguiar Dos Santos1, Álvaro Farias Pinheiro2, Adelino Ferreira1
1CITTA - Research Centre for Territory, Transport and Environment, Department of Civil Engineering, University of Coimbra, Coimbra, Portugal.
Objective:
This study aims to develop an Explainable Artificial Intelligence (XAI) framework to classify and predict road accident severity on a strategic road network in a developing country. The research focuses on identifying the primary risk factors contributing to fatal and serious injuries to support "Vision Zero" policies.
Methods:
Three machine learning algorithms-Random Forest (RF), XGBoost, and Artificial Neural Networks (ANN)-were trained using a dataset of traffic accidents from federal highways in Pernambuco, Brazil. The study specifically addressed data imbalance by comparing model performance on the original data distribution versus the Synthetic Minority Over-sampling Technique (SMOTE). To ensure transparency, the SHapley Additive exPlanations (SHAP) technique was applied to interpret the models' decision-making process.
Results:
The Random Forest model, when applied to the original data distribution, demonstrated superior performance and methodological rigor, achieving an AUC-ROC of 0.952 and an F1-Score of 0.747. The SHAP analysis revealed that accident lethality is primarily driven by the synergy between single-lane roads, specific collision types (such as head-on collisions), and excessive speed.
Conclusions:
The findings suggest that preserving the latent properties of original data, rather than artificial balancing, provides a more reliable foundation for safety interventions. The integration of XAI tools allowed for the decoding of complex "black box" models, providing actionable insights for transportation engineers and policymakers to prioritize infrastructure improvements and enforcement in emerging economies.