XGBoost-Based Feature Learning Method for Mining COVID-19 Novel Diagnostic Markers

Xianbin Song1, Jiangang Zhu1, Xiaoli Tan2

  • 1Department of Critical Care Medicine, Affiliated Hospital of Jiaxing University, Jiaxing, China.

Insights

Researchers identified 24 novel gene markers for diagnosing COVID-19. These genes effectively distinguish between positive and negative patients, offering potential for new diagnostic tools.

Area of Science:

  • Genomics
  • Infectious Diseases
  • Bioinformatics

Background:

  • The COVID-19 pandemic, caused by SARS-CoV-2, emerged in late 2019, leading to a global health crisis and significant economic impact.
  • Accurate and rapid diagnostic methods are crucial for managing infectious disease outbreaks.

Purpose of the Study:

  • To identify novel diagnostic biomarkers for COVID-19 using gene expression data.
  • To develop and validate machine learning models for classifying COVID-19 patients.

Main Methods:

  • Downloaded and analyzed throat swab gene expression data from COVID-19 positive and negative patients via the Gene Expression Omnibus (GEO) database.
  • Employed XGBoost for feature gene selection and constructed various machine learning classifiers (MARS, KNN, SVM, MIL, RF).
  • Utilized the Iterative Feature Selection (IFS) method to select the optimal KNN classifier and identified 24 feature genes, validated using Principal Component Analysis (PCA).

Main Results:

  • Identified a set of 24 feature genes capable of effectively classifying COVID-19 positive and negative patients.
  • The selected genes were significantly enriched in biological functions related to viral transcription and viral gene expression.
  • Pathway analysis indicated enrichment in pathways associated with Coronavirus disease-COVID-19.

Conclusions:

  • The 24 identified feature genes demonstrate high efficacy in distinguishing between COVID-19 positive and negative individuals.
  • These genes hold promise as novel biomarkers for the diagnosis of COVID-19.
  • The findings contribute to the development of more effective diagnostic strategies for the pandemic.

Related Concept Videos

Classification of Illness01:17

Classification of Illness

The meaning of illness is individualized to each person who experiences an alteration in health. In contrast, disease is a medical term indicating a pathological change in the structure and function of the body or mind. It is a condition that has specific symptoms and boundaries.
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
7.9K
Sensitivity, Specificity, and Predicted Value01:13

Sensitivity, Specificity, and Predicted Value

In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
636
Cancer Survival Analysis01:21

Cancer Survival Analysis

Cancer survival analysis focuses on quantifying and interpreting the time from a key starting point, such as diagnosis or the initiation of treatment, to a specific endpoint, such as remission or death. This analysis provides critical insights into treatment effectiveness and factors that influence patient outcomes, helping to shape clinical decisions and guide prognostic evaluations. A cornerstone of oncology research, survival analysis tackles the challenges of skewed, non-normally...
442
Single Nucleotide Polymorphisms-SNPs01:05

Single Nucleotide Polymorphisms-SNPs

A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
15.7K
Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.9K
Aggregates Classification01:29

Aggregates Classification

Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
373