Related Experiment Video
Updated: Jan 9, 2026

Trajectory Data Analyses for Pedestrian Space-time Activity Study
Published on: February 25, 2013
Spatial prediction of COVID-19 pandemic dynamics in the United States
Çiğdem Ak1, Alex D Chitsazan1, Mehmet Gönen2
1Cancer Early Detection Advanced Research Center, Knight Cancer Institute, Oregon Health & Science University, 2720 S Moody Ave, Portland, OR 97201, USA.
Abstract:
The impact of COVID-19 across the United States has been heterogeneous, with rapid spread and greater mortality in some areas compared with others. We used geographically-linked data to test the hypothesis that the risk for COVID-19 is defined by location and sought to define which demographic features are most closely associated with elevated COVID-19 spread and mortality. We leveraged geographically-restricted social, economic, political, and demographic information from US counties, to develop a computational framework using structured Gaussian processing to predict county-level case and death counts during the pandemic's initial and nationwide phases. After identifying the most predictive information sources by location, we applied an unsupervised clustering algorithm and topic modelling to identify groups of features most closely associated with COVID-19 spread. Our model successfully predicted COVID-19 case counts of unseen locations, after examining case counts and demographic information of neighboring locations, with overall Pearson's correlation coefficient and the proportion of variance explained of 0.96 and 0.84 during the initial phase and 0.95 and 0.87, respectively, during the nationwide phase. Aside from population metrics, presidential vote margin was the most consistently selected spatial feature in our COVID-19 prediction models. Urbanicity and 2020 presidential vote margins were more predictive than other demographic features. Models trained using death counts showed similar performance metrics. Topic modeling showed that counties with similar socioeconomic and demographic features tended to group together, and some of these grouped feature sets were associated with COVID-19 dynamics. Clustering of counties based on these feature groups found by topic modeling revealed groups of counties that experienced markedly different COVID-19 spread. We conclude that topic modeling can be used to group similar features and identify counties with similar features in epidemiologic research.
Related Concept Videos
Steps in Outbreak Investigation
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Causality in Epidemiology
Pareto Chart
The Pareto chart is named after the Italian economist Vilfredo Pareto, who described the Pareto...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Statistical Methods for Analyzing Epidemiological Data

