Supervised Machine Learning Models for Prediction of COVID-19 Infection using Epidemiology Dataset
L J Muhammad1, Ebrahem A Algehyne2, Sani Sharif Usman3
1Department of Mathematics and Computer Science, Faculty of Science, Federal University of Kashere, P.M.B. 0182, Gombe, Nigeria.
Summary
Machine learning models show promise for diagnosing COVID-19 (2019-nCoV) when specific treatments are unavailable. The decision tree model achieved 94.99% accuracy, aiding prognosis and reducing healthcare burdens.
Area of Science:
- Medical Informatics
- Computational Biology
- Epidemiology
Background:
- COVID-19 (2019-nCoV) is now endemic, posing an ongoing challenge to global healthcare systems due to the lack of specific treatments or cures.
- The absence of effective antivirals or vaccines necessitates alternative strategies to manage the disease and alleviate healthcare burdens, particularly in developing nations.
Purpose of the Study:
- To develop and evaluate supervised machine learning models for diagnosing COVID-19 (2019-nCoV) infection.
- To assess the performance of various machine learning algorithms in classifying COVID-19 cases using epidemiological data.
Main Methods:
- Employed supervised machine learning algorithms including logistic regression, decision tree, support vector machine, naive Bayes, and artificial neural network.
- Utilized an epidemiology-labeled dataset of positive and negative COVID-19 cases from Mexico, with 80% for training and 20% for testing.
- Performed correlation coefficient analysis to understand feature relationships before model development.
Main Results:
- The decision tree model demonstrated the highest accuracy at 94.99%.
- The Support Vector Machine model achieved the highest sensitivity (93.34%).
- The Naïve Bayes model exhibited the highest specificity (94.30%).
Conclusions:
- Supervised machine learning models offer a viable approach for COVID-19 (2019-nCoV) diagnosis and prognosis.
- These AI-driven methods can help reduce the strain on healthcare systems and economic sectors impacted by the endemic disease.
Related Concept Videos
Steps in Outbreak Investigation
369
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
369
Statistical Methods for Analyzing Epidemiological Data
740
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
740
Introduction to Epidemiology
1.4K
Epidemiology, known as the cornerstone of public health, involves studying the distribution and determinants of health-related events in defined populations and applying these insights to control health issues. This is essential for understanding how diseases spread, identifying populations at greater risk, and implementing measures to control or prevent outbreaks. Epidemiology addresses not only infectious diseases but also non-communicable conditions like cancer and cardiovascular disease,...
1.4K
Statistical Software for Data Analysis and Clinical Trials
1.2K
Statistical software is pivotal in data analysis and clinical trials by providing tools to analyze data, draw conclusions, and make predictions. These software packages range from simple data management applications to complex analytical platforms, supporting various statistical tests, models, and simulation techniques. Their significance lies in their ability to handle vast amounts of data with precision and efficiency, enabling researchers to validate hypotheses, identify trends, and make...
1.2K
Classification of Illness
8.3K
The meaning of illness is individualized to each person who experiences an alteration in health. In contrast, disease is a medical term indicating a pathological change in the structure and function of the body or mind. It is a condition that has specific symptoms and boundaries.
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
8.3K
Residuals and Least-Squares Property
8.6K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
8.6K


