An Improved Machine-Learning Approach for COVID-19 Prediction Using Harris Hawks Optimization and Feature Analysis
Kumar Debjit1, Md Saiful Islam2, Md Abadur Rahman3
1Faculty of Health, Engineering and Sciences, University of Southern Queensland, 487-535 West Street, Toowoomba, QLD 4350, Australia.
Diagnostics (Basel, Switzerland)
|May 28, 2022
Summary
This study introduces an optimized machine learning (ML) framework using Harris Hawks Optimization (HHO) for early COVID-19 detection. The ensemble model achieved 92.38% accuracy, outperforming traditional methods.
Area of Science:
- Computational biology
- Medical informatics
- Epidemiology
Background:
- The COVID-19 pandemic highlighted the need for advanced healthcare monitoring systems.
- Artificial intelligence (AI) and machine learning (ML) are crucial for analyzing big data generated during pandemics.
- Early detection of COVID-19 is vital for effective disease control and patient management.
Purpose of the Study:
- To propose an improved ML framework for the early detection of COVID-19.
- To optimize ML hyperparameters using the Harris Hawks Optimization (HHO) algorithm.
- To enhance prediction performance through an ensemble technique.
Main Methods:
- Applied HHO algorithm to optimize hyperparameters of ML models: eXtreme gradient boosting (XGBoost), light gradient boosting, categorical boosting, random forest, and support vector classifier.
- Utilized an ensemble technique combining optimized ML models for improved prediction.
- Evaluated feature importance using SHapely adaptive exPlanations (SHAP) values.
Main Results:
- The proposed ensemble ML model achieved a prediction accuracy of 92.38% on publicly available COVID-19 big data.
- The HHO-optimized eXtreme gradient boosting (HHOXGB) model showed the highest single-model accuracy at 92.23%.
- The proposed method demonstrated superior performance compared to traditional and other ML-based approaches.
Conclusions:
- The developed ML framework effectively improves early COVID-19 detection accuracy.
- Feature importance analysis provides insights into key indicators for disease detection.
- A graphical user interface is proposed for accessibility by non-specialist healthcare professionals.
Related Concept Videos
Steps in Outbreak Investigation
221
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
221
Prediction Intervals
2.4K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.4K
Statistical Methods for Analyzing Epidemiological Data
566
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
566
Residuals and Least-Squares Property
7.9K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.9K
Improving Translational Accuracy
11.9K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.9K
Statistical Software for Data Analysis and Clinical Trials
812
Statistical software is pivotal in data analysis and clinical trials by providing tools to analyze data, draw conclusions, and make predictions. These software packages range from simple data management applications to complex analytical platforms, supporting various statistical tests, models, and simulation techniques. Their significance lies in their ability to handle vast amounts of data with precision and efficiency, enabling researchers to validate hypotheses, identify trends, and make...
812


