Related Experiment Video
Updated: Oct 25, 2025

Data Acquisition Protocol for Determining Embedded Sensitivity Functions
Published on: April 20, 2016
Examining the impacts of crash data aggregation on SPF estimation
Agnimitra Sengupta1, Vikash V Gayah1, Eric T Donnell1
1Department of Civil and Environmental Engineering, The Pennsylvania State University, 231 Sackett Building, University Park, PA 16802, United States.
Abstract:
The American Association of State Highway and Transportation Officials' Highway Safety Manual (HSM) includes a collection of safety performance functions (SPFs) or statistical models to estimate the expected crash frequency of roadway segments, intersections, and interchanges. These models are applied in several steps of the safety management process, including to screen the road network for opportunities to improve safety and to evaluate the performance of safety countermeasure deployments. The SPFs in the HSM are generally estimated using negative binomial regression modeling. In some instances, they are estimated using annual crash frequency and site-specific (e.g., traffic volume) data, while in other instances they are estimated using aggregate crash frequency and site-specific data. This paper explores the differences that result from estimating SPFs using aggregate versus disaggregate data using the same methods as those used to estimate the SPFs in the HSM. A synthetic dataset was first used to conduct these comparisons - these data were generated in a manner that is consistent with the properties of the negative binomial distribution. Then, an observational dataset from Pennsylvania was used to compare the SPFs from both aggregate and disaggregate data. The results show that SPFs estimated using the panel (disaggregate) data and aggregated data provide similar model coefficients, although some differences may sometimes arise. However, the overdispersion parameter obtained using each dataset can differ significantly. These differences result in systematic biases in calculations of expected crash frequency when Empirical Bayes adjustments are applied, which - as the paper demonstrates - could lead to different outcomes in a network screening exercise. Overall, these results reveal that aggregating crash data might result in biased SPF outputs and lead to inconsistent Empirical Bayes adjustments.
Related Concept Videos
Determination of Expected Frequency
Maximum Size of Aggregate
Elastic Collisions: Case Study
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Unsoundness of Aggregate due to Volume Change
Hypothesis Test for Test of Independence
H0: The two variables (factors)...

