Related Experiment Video
Updated: Oct 30, 2025

Author Spotlight: Advancements in Multiplex Detection of Respiratory Viruses
Published on: November 10, 2023
Machine Learning for Analyzing Non-Countermeasure Factors Affecting Early Spread of COVID-19
Vito Janko1, Gašper Slapničar1, Erik Dovgan1
1Jožef Stefan Institute, 1000 Ljubljana, Slovenia.
Abstract:
The COVID-19 pandemic affected the whole world, but not all countries were impacted equally. This opens the question of what factors can explain the initial faster spread in some countries compared to others. Many such factors are overshadowed by the effect of the countermeasures, so we studied the early phases of the infection when countermeasures had not yet taken place. We collected the most diverse dataset of potentially relevant factors and infection metrics to date for this task. Using it, we show the importance of different factors and factor categories as determined by both statistical methods and machine learning (ML) feature selection (FS) approaches. Factors related to culture (e.g., individualism, openness), development, and travel proved the most important. A more thorough factor analysis was then made using a novel rule discovery algorithm. We also show how interconnected these factors are and caution against relying on ML analysis in isolation. Importantly, we explore potential pitfalls found in the methodology of similar work and demonstrate their impact on COVID-19 data analysis. Our best models using the decision tree classifier can predict the infection class with roughly 80% accuracy.
More Related Videos
Related Concept Videos
Steps in Outbreak Investigation
Factors Affecting the Risk of Infection
The integrity and count of the white blood cells help the body resist pathogens and fight infection. When impaired, it reduces the body's resistance to pathogens. The acidic pH levels of the gastrointestinal, genitourinary tracts, and skin...
Factors Affecting Illness
For instance, risk factors are connected to illness,...
Statistical Methods for Analyzing Epidemiological Data
Causality in Epidemiology
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...

