A novel computational method for assigning weights of importance to symptoms of COVID-19 patients

Mohammad A Alzubaidi1, Mwaffaq Otoom1, Nesreen Otoum2

  • 1Department of Computer Engineering, Yarmouk University, Irbid, 21163, Jordan.

Insights

Identifying key symptoms of coronavirus disease 2019 (COVID-19) is crucial. This study found that fever, cough, fatigue, sore throat, and shortness of breath are the most significant indicators of COVID-19.

Area of Science:

  • Medical Informatics
  • Epidemiology
  • Data Science

Background:

  • The COVID-19 pandemic presents a challenge with diverse patient symptoms.
  • Understanding common symptoms and their importance is critical for diagnosis and management.

Purpose of the Study:

  • To identify and rank the most important symptoms of COVID-19.
  • To evaluate the effectiveness of feature selection algorithms in analyzing COVID-19 symptom data.

Main Methods:

  • A COVID-19 dataset of 738 confirmed cases was preprocessed.
  • Six feature selection algorithms, including a novel Variance Based Feature Weighting (VBFW) method, were applied.
  • Symptom importance was ranked and quantitatively measured.

Main Results:

  • Aggregated results from five algorithms identified Fever/Cough, Fatigue, Sore Throat, and Shortness of Breath as key symptoms.
  • The novel VBFW algorithm ranked Fever (75%) and Cough (39.8%) as highly indicative.
  • VBFW achieved 92.1% accuracy in a one-class SVM model and 100% NDCG@5.

Conclusions:

  • Fever, Cough, Fatigue, Sore Throat, and Shortness of Breath are important indicators of COVID-19.
  • The VBFW algorithm highlights Fever and Cough as particularly significant symptoms in the analyzed dataset.
Abstract

Related Concept Videos

Weighted Mean00:57

Weighted Mean

While taking the arithmetic, geometric, or harmonic mean of a sample data set, equal importance is assigned to all the data points. However, all the values may not always be equally important in some data sets. An intrinsic bias might make it more important to give more weightage to specific values over others.
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
6.0K
Classification of Illness01:17

Classification of Illness

The meaning of illness is individualized to each person who experiences an alteration in health. In contrast, disease is a medical term indicating a pathological change in the structure and function of the body or mind. It is a condition that has specific symptoms and boundaries.
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
8.2K
Pareto Chart00:52

Pareto Chart

A Pareto chart is a bar graph or a combination of both line and bar graphs. The bar lengths represent the individual values or the frequency, while the lines represent the cumulative total values. In this chart, the longest bars are arranged on the left and the shortest bars on the right, which makes it easier to read and interpret the data. It can also be called a Pareto diagram or Pareto analysis.
The Pareto chart is named after the Italian economist Vilfredo Pareto, who described the Pareto...
7.4K
Statistical Methods for Analyzing Epidemiological Data01:25

Statistical Methods for Analyzing Epidemiological Data

Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
712
Residuals and Least-Squares Property01:11

Residuals and Least-Squares Property

The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
8.4K