Related Experiment Video
Updated: Jul 31, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Analyzing Impact of Socio-Economic Factors on COVID-19 Mortality Prediction Using SHAP Value
Redoan Rahman1, Jooyeong Kang1, Justin F Rousseau2
1School of Information, The University of Texas at Austin, Austin, Texas, United States.
Abstract:
This paper applies multiple machine learning (ML) algorithms to a dataset of de-identified COVID-19 patients provided by the COVID-19 Research Database. The dataset consists of 20,878 COVID-positive patients, among which 9,177 patients died in the year 2020. This paper aims to understand and interpret the association of socio-economic characteristics of patients with their mortality instead of maximizing prediction accuracy. According to our analysis, a patient's household's annual and disposable income, age, education, and employment status significantly impacts a machine learning model's prediction. We also observe several individual patient data, which gives us insight into how the feature values impact the prediction for that data point. This paper analyzes the global and local interpretation of machine learning models on socio-economic data of COVID patients.
More Related Videos
08:51Author Spotlight: Integrated Multi-Omics Analysis for Unveiling Multicellular Immune Signatures in Clinical Heart Attack Cohorts
Published on: September 20, 2024
05:15Cutoff Value of Phase Angle by Bioelectrical Impedance Analysis at Admission as a Prognostic Factor in Patients with Acute Heart Failure
Published on: June 10, 2025
Related Concept Videos
Bias in Epidemiological Studies
Statistical Methods for Analyzing Epidemiological Data
Factors Affecting Illness
For instance, risk factors are connected to illness,...
Causality in Epidemiology
Applications of Life Tables
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...