Related Experiment Video
Updated: Jun 25, 2025

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Interpretable machine learning for evaluating risk factors of freeway crash severity
Seyed Alireza Samerei1, Kayvan Aghabayk1
1School of Civil Engineering, College of Engineering, University of Tehran, Tehran, Iran.
Abstract:
Machine learning (ML) models are widely employed for crash severity modelling, yet their interpretability remains underexplored. Interpretation is crucial for comprehending ML results and aiding informed decision-making. This study aims to implement an interpretable ML to visualize the impacts of factors on crash severity using 5 years of freeways data from Iran. Methods including classification and regression trees (CART), K-nearest neighbours (KNNs), random forest (RF), artificial neural network (ANN) and support vector machines (SVM) were applied, with RF demonstrating superior accuracy, recall, F1-score and ROC. The accumulated local effects (ALE) were utilized for interpretation. Findings suggest that light traffic conditions () with critical values around 0.05 or 0.38, and higher proportion of large trucks and buses, particularly at 10% and 4%, are associated with severe crashes. Additionally, speeds exceeding 90 km/h, drivers younger than 30 years, rollover crashes, collisions with fixed objects and barriers, nighttime driving and driver fatigue elevate the likelihood of severe crashes.
More Related Videos
Related Concept Videos
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Determination of Expected Frequency
Relative Risk
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Hypothesis Test for Test of Independence
H0: The two variables (factors)...
Statistical Methods for Analyzing Epidemiological Data

