Related Experiment Video
Updated: Sep 1, 2025

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Supervised Pretraining through Contrastive Categorical Positive Samplings to Improve COVID-19 Mortality Prediction
Tingyi Wanyan1, Mingquan Lin1, Eyal Klang2
1Population Health Sciences, Weill Cornell Medicine, New York, NY, USA.
Abstract:
Clinical EHR data is naturally heterogeneous, where it contains abundant sub-phenotype. Such diversity creates challenges for outcome prediction using a machine learning model since it leads to high intra-class variance. To address this issue, we propose a supervised pre-training model with a unique embedded k-nearest-neighbor positive sampling strategy. We demonstrate the enhanced performance value of this framework theoretically and show that it yields highly competitive experimental results in predicting patient mortality in real-world COVID-19 EHR data with a total of over 7,000 patients admitted to a large, urban health system. Our method achieves a better AUROC prediction score of 0.872, which outperforms the alternative pre-training models and traditional machine learning methods. Additionally, our method performs much better when the training data size is small (345 training instances).
Related Concept Videos
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Improving Translational Accuracy
Comparing the Survival Analysis of Two or More Groups
Cancer Survival Analysis
Survival Tree
Building a Survival Tree
Constructing a...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.

