A descriptive study of random forest algorithm for predicting COVID-19 patients outcome
Jie Wang1, Heping Yu2, Qingquan Hua1
1Department of Otolaryngology-Head and Neck Surgery, Renmin Hospital of Wuhan University, Wuhan, Hubei, China.
Insights
This study used a random forest algorithm to predict COVID-19 patient mortality. High levels of lactate dehydrogenase (LDH) and myoglobin (Myo) were identified as key predictors of poor prognosis in COVID-19 patients.
Area of Science:
- Medical research
- Clinical diagnostics
- Public health
Background:
- The COVID-19 pandemic poses a significant global health challenge.
- Identifying reliable predictors for COVID-19 patient outcomes is crucial for effective clinical management.
Purpose of the Study:
- To develop a predictive model for COVID-19 patient prognosis using machine learning.
- To identify key clinical indicators for predicting mortality in COVID-19 patients.
Main Methods:
- Collected clinical data from 126 COVID-19 patients.
- Applied a random forest (RF) algorithm for prognosis prediction.
- Utilized SMOTE and RFE for data balancing and feature selection.
Main Results:
- The RF model achieved 100% accuracy in predicting COVID-19 patient prognoses.
- Lactate dehydrogenase (LDH) and Myoglobin (Myo) were identified as optimal predictors.
- Elevated LDH (>500 U/L) and Myo (>80 ng/ml) significantly increased mortality risk.
Conclusions:
- An RF algorithm accurately predicts COVID-19 patient mortality.
- LDH and Myo levels are valuable early indicators for assessing COVID-19 patient prognosis.
- These findings can aid in early risk stratification and clinical decision-making for COVID-19 patients.
Background:
The outbreak of coronavirus disease 2019 (COVID-19) that occurred in Wuhan, China, has become a global public health threat. It is necessary to identify indicators that can be used as optimal predictors for clinical outcomes of COVID-19 patients.
Methods:
The clinical information from 126 patients diagnosed with COVID-19 were collected from Wuhan Fourth Hospital. Specific clinical characteristics, laboratory findings, treatments and clinical outcomes were analyzed from patients hospitalized for treatment from 1 February to 15 March 2020, and subsequently died or were discharged. A random forest (RF) algorithm was used to predict the prognoses of COVID-19 patients and identify the optimal diagnostic predictors for patients' clinical prognoses.
Results:
Seven of the 126 patients were excluded for losing endpoints, 103 of the remaining 119 patients were discharged (alive) and 16 died in the hospital. A synthetic minority over-sampling technique (SMOTE) was used to correct the imbalanced distribution of clinical patients. Recursive feature elimination (RFE) was used to select the optimal subset for analysis. Eleven clinical parameters, Myo, CD8, age, LDH, LMR, CD45, Th/Ts, dyspnea, NLR, D-Dimer and CK were chosen with AUC approximately 0.9905. The RF algorithm was built to predict the prognoses of COVID-19 patients based on the best subset, and the area under the ROC curve (AUC) of the test data was 100%. Moreover, two optimal clinical risk predictors, lactate dehydrogenase (LDH) and Myoglobin (Myo), were selected based on the Gini index. The univariable logistic analysis revealed a substantial increase in the risk for in-hospital mortality when Myo was higher than 80 ng/ml (OR = 7.54, 95% CI [3.42-16.63]) and LDH was higher than 500 U/L (OR = 4.90, 95% CI [2.13-11.25]).
Conclusion:
We applied an RF algorithm to predict the mortality of COVID-19 patients with high accuracy and identified LDH higher than 500 U/L and Myo higher than 80 ng/ml to be potential risk factors for the prognoses of COVID-19 patients in the early stage of the disease.
More Related Videos
Related Concept Videos
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Statistical Methods for Analyzing Epidemiological Data
Survival Tree
Building a Survival Tree
Constructing a...
Contingency Table
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Study Design in Statistics
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...


