Related Experiment Video
Updated: Jul 18, 2026

16:14
Trajectory Data Analyses for Pedestrian Space-time Activity Study
Published on: February 25, 2013
Application of hidden Markov models on residuals: an example using Canadian traffic accident data
W H Laverty1, M J Miket, I W Kelly
1Department of Educational Psychology and Special Education, University of Saskatchewan, Saskatoon, Canada.
Perceptual and Motor Skills
|August 21, 2002
Summary
Hidden Markov models revealed two distinct accident patterns in Saskatchewan automobile data. A
Area of Science:
- Social Science
- Statistics
- Transportation Safety
Background:
- Road safety analysis often involves identifying trends and seasonal variations in accident data.
- Previous analyses of Saskatchewan automobile accidents (1992) identified linear trends, seasonal, holiday, and day-of-week effects.
Purpose of the Study:
- To explore underlying patterns in automobile accident data beyond traditional regression analysis.
- To apply hidden Markov models (HMMs) to residual data from a regression analysis of traffic accidents.
Main Methods:
- Conducted a regression analysis on 9 years of Saskatchewan automobile accident data (200,545 accidents).
- Applied a hidden Markov model to the residuals of the regression analysis.
- Identified and characterized distinct states within the residual data.
Main Results:
- Uncovered two hidden states in the accident data, labeled 'low volatility' and 'high volatility'.
- The 'high volatility' state, characterized by greater variability, was associated with colder months and a higher number of accidents.
- The 'low volatility' state exhibited less variability.
Conclusions:
- Hidden Markov models are effective for detecting unobserved states in complex datasets.
- The identified states likely relate to weather conditions influencing accident frequency and severity.
- HMMs offer a valuable tool for uncovering hidden dynamics in social science and health-related research.
Related Concept Videos
Probability Histograms
A probability histogram is a visual representation of a probability distribution. Similar a typical histogram, the probability histogram consists of contiguous (adjoining) boxes. It has both a horizontal axis and a vertical axis. The horizontal axis is labeled with what the data represents. The vertical axis is labeled with probability. Each rectangular bar in the histogram is 1 unit wide, which suggests that the area under each bar equals the probability, P(x), where x is 1, 2, 3, and so on.
Hypothesis Test for Test of Independence
The test of independence is a chi-square-based test used to determine whether two variables or factors are independent or dependent. This hypothesis test is used to examine the independence of the variables. One can construct two qualitative survey questions or experiments based on the variables in a contingency table. The goal is to see if the two variables are unrelated (independent) or related (dependent). The null and alternative hypotheses for this test are:
H0: The two variables (factors)...
H0: The two variables (factors)...
Determination of Expected Frequency
Suppose one wants to test independence between the two variables of a contingency table. The values in the table constitute the observed frequencies of the dataset. But how does one determine the expected frequency of the dataset? One of the important assumptions is that the two variables are independent, which means the variables do not influence each other. For independent variables, the statistical probability of any event involving both variables is calculated by multiplying the individual...
Residuals and Least-Squares Property
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Residual Plots
A residual plot is a statistical representation of data used to analyze correlation and regression results. It helps verify the requirements for drawing specific conclusions about correlation and regression. To obtain the residual plot, first, the residual for each data value is calculated, which is simply the vertical distance between the observed and the predicted value obtained from the regression equation.
When the residual values are plotted against the variable x, it is called a residual...
When the residual values are plotted against the variable x, it is called a residual...
Mechanistic Models: Compartment Models in Individual and Population Analysis
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least squares (OLS)...
