Related Experiment Video
Updated: Jun 26, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Maximally informative feature selection using Information Imbalance: Application to COVID-19 severity prediction
Romina Wild1, Emanuela Sozio2,3, Riccardo G Margiotta1
1International School for Advanced Studies (SISSA), Via Bonomea 265, 34136, Trieste, Italy.
This study introduces a new statistical method to identify key patient features for predicting COVID-19 outcomes. The approach effectively selects informative features from complex clinical data, aiding in disease severity assessment.
Area of Science:
- Biostatistics
- Clinical Informatics
- Data Science
Background:
- Clinical databases contain diverse patient data, including medical history, symptoms, and test results.
- Identifying the most informative features is crucial for accurate clinical predictions, especially for minority patient groups.
Purpose of the Study:
- To adapt and apply the Information Imbalance statistical approach for selecting maximally informative patient features in a clinical setting.
- To identify key features predictive of clinical fate and disease severity in COVID-19 patients.
Main Methods:
- Utilized the Information Imbalance statistical approach, adapted for categorical and incomplete clinical data.
- Applied the algorithm to a dataset of approximately 1300 COVID-19 patients treated before October 2021.
- The method automatically determines the optimal number of features and handles missing data without imputation.
Main Results:
- Identified combinations of 10-15 patient features, measurable at admission, that are highly informative of clinical outcomes and disease severity.
- Demonstrated the approach's effectiveness even with features available for only a fraction of patients.
- The selected features exhibited low inter-feature correlation.
Conclusions:
- The adapted Information Imbalance method provides an effective way to select parsimonious, informative feature sets from complex clinical data.
- This approach can enhance predictive modeling for diseases like COVID-19, offering valuable clinical insights for patient management.
Related Concept Videos
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Bias in Epidemiological Studies
Statistical Methods for Analyzing Epidemiological Data
Quantifying and Rejecting Outliers: The Grubbs Test
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...

