Related Experiment Videos
Genetic determinants of metabolic-inflammatory dysregulation and machine learning prediction of COVID-19
Larysa Sydorchuk1, Maksym Sokolenko2, Ruslan Sydorchuk3
1Department of Family Medicine, Bukovinian State Medical University, Chernivtsi, Ukraine.
Abstract:
The aim of this study was to identify genetic variations that influence the metabolic-inflammatory profile of coronavirus disease and to evaluate the performance of modern machine learning models for the classification and prediction of COVID-19 severity. Real-time polymerase chain reaction was used to genotype the polymorphism of FGB (rs1800790), NOS3 (rs2070744) and TMPRSS2 (rs12329760) genes. Model performance was assessed using Accuracy and AUC-ROC, with a focus on maximizing AUC-ROC to ensure optimal discrimination, and SHAP analysis. Optimal hyperparameters for each model were defined as those yielding the highest mean AUC-ROC value during 5-fold cross-validation. COVID-19 severity is strongly associated with a pronounced pro-inflammatory and pro-endothelial activation profile, characterized by significantly elevated transmembrane serine protease 2 (TMPRSS2), endothelin-1 (ET-1), interleukin-6 (IL-6) and procalcitonin (PCT) levels, together with marked metabolic dysregulation. Genetic variations contribute to inter-individual differences in biomarker expression, with specific allelic variants modulating inflammatory intensity and endothelial dysfunction. In particular, the FGB rs1800790 A-allele and eNOS rs2070744 ТТ-genotype are associated with a more pronounced inflammatory response and endothelial dysfunction, while rs12329760 TMPRSS2 variants show weaker but detectable modulatory effects with higher transmembrane serine protease 2 value in T-allele moderate-severe COVID-19 patients. Ensemble methods such as ExtraTreesClassifier (Accuracy: 0.974 ± 0.022) and RandomForestClassifier (Accuracy: 0.960 ± 0.035) demonstrated the highest ROC curves approaching, confirming their encouraging performance in predicting COVID-19 severity. Simpler models, including BernoulliNB (Accuracy: 0.956 ± 0.037) and DecisionTreeClassifier (Accuracy: 0.938 ± 0.043), also showed high classification quality. Analysis of misclassification patterns revealed that ExtraTreesClassifier, HistGradientBoostingClassifier, BaggingClassifier, and GradientBoostingClassifier made no errors across any class. The poorest performance was observed with LinearDiscriminantAnalysis, which generated 11 misclassifications, followed by CalibratedClassifierCV, and LogisticRegressionCV. External validation of the obtained results in larger multicentre cohorts is essential before clinical implementation.
Related Concept Videos
Human Genetics
The complex relationship between genetics and psychology is observable through common biological components such...
Type II Diabetes I: Introduction
Pharmacogenomics: Identification of New Drug Targets