Related Experiment Video
Updated: Jun 25, 2025

12:18
A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
7.5K
Interpretable machine learning for evaluating risk factors of freeway crash severity
Seyed Alireza Samerei1, Kayvan Aghabayk1
1School of Civil Engineering, College of Engineering, University of Tehran, Tehran, Iran.
Summary
Interpretable machine learning reveals factors influencing crash severity. Light traffic, large trucks, high speeds, young drivers, and nighttime driving increase severe crash risk on Iranian freeways.
Area of Science:
- Traffic Safety Engineering
- Data Science
- Transportation Research
Background:
- Machine learning (ML) models are increasingly used for crash severity prediction.
- Interpretability of these ML models is often overlooked, limiting practical insights.
- Understanding factors contributing to crash severity is vital for effective safety interventions.
Purpose of the Study:
- To implement interpretable machine learning techniques for visualizing factors affecting crash severity.
- To analyze five years of freeway crash data from Iran to identify key risk factors.
- To enhance decision-making processes in traffic safety through clear model interpretation.
Main Methods:
- Applied various machine learning models: Classification and Regression Trees (CART), K-nearest neighbours (KNNs), Random Forest (RF), Artificial Neural Network (ANN), and Support Vector Machines (SVM).
- Utilized Accumulated Local Effects (ALE) plots for model interpretation.
- Evaluated model performance using accuracy, recall, F1-score, and ROC metrics, with RF showing superior results.
Main Results:
- Random Forest model demonstrated the highest performance in predicting crash severity.
- Key factors associated with severe crashes include light traffic conditions (critical values around 0.05 or 0.38).
- A higher proportion of large trucks and buses (e.g., 10% and 4%), speeds over 90 km/h, drivers under 30, rollover crashes, collisions with fixed objects/barriers, nighttime driving, and driver fatigue significantly increase severe crash likelihood.
Conclusions:
- Interpretable ML, specifically RF with ALE, effectively visualizes and quantifies the impact of various factors on crash severity.
- Findings provide actionable insights for targeted safety measures on freeways.
- The study highlights the importance of driver demographics, vehicle types, environmental conditions, and crash dynamics in determining severity.
More Related Videos
Related Concept Videos
Regression Analysis
5.7K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.7K
Determination of Expected Frequency
2.2K
Suppose one wants to test independence between the two variables of a contingency table. The values in the table constitute the observed frequencies of the dataset. But how does one determine the expected frequency of the dataset? One of the important assumptions is that the two variables are independent, which means the variables do not influence each other. For independent variables, the statistical probability of any event involving both variables is calculated by multiplying the individual...
2.2K
Relative Risk
151
Relative risk (RR) is a statistical measure commonly used in epidemiology to compare the likelihood of a particular event occurring between two groups. This metric is important for evaluating the relationship between exposure to a specific risk factor and the probability of a particular outcome. It plays a crucial role in medical research, public health studies, and risk assessment. Relative risk quantifies how much more (or less) likely an event is to occur in an exposed group compared to an...
151
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Hypothesis Test for Test of Independence
3.6K
The test of independence is a chi-square-based test used to determine whether two variables or factors are independent or dependent. This hypothesis test is used to examine the independence of the variables. One can construct two qualitative survey questions or experiments based on the variables in a contingency table. The goal is to see if the two variables are unrelated (independent) or related (dependent). The null and alternative hypotheses for this test are:
H0: The two variables (factors)...
H0: The two variables (factors)...
3.6K
Statistical Methods for Analyzing Epidemiological Data
361
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
361

