Related Experiment Video
Updated: Apr 19, 2026

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Handling varying amounts of missing data when classifying mental-health risk levels
Sherine Nagy Saleh1, Christopher D Buckingham2
1Arab Academy for Science and Technology, Alexandria, Egypt.
This study introduces a novel method for handling missing clinical data, improving accuracy in predictive models. The approach ensures risk explanations are based on available data, crucial for clinical decision support systems.
Area of Science:
- Clinical Data Science
- Machine Learning in Healthcare
- Predictive Analytics
Background:
- Classifying clinical data presents challenges due to missing features.
- Traditional methods like imputation or data deletion can reduce accuracy with high missingness.
- Existing approaches struggle with datasets having substantial missing values and high variability.
Purpose of the Study:
- To propose a new methodology for handling clinical datasets with a high percentage of missing values.
- To develop a classification approach that accommodates high variability in missing data patterns.
- To improve the accuracy and interpretability of predictive models in clinical settings.
Main Methods:
- Sequential feature selection based on correlation with the dependent variable and selected features.
- Individual classification model generation for each test case using available training data.
- Application to real-world mental health data for suicide risk prediction.
Main Results:
- The proposed method demonstrates favorable comparison with alternative approaches.
- The methodology effectively handles datasets with significant missing data and high variability.
- Explanations for risk predictions are derived solely from provided data, not imputed values.
Conclusions:
- The new methodology offers an effective solution for classifying clinical data with extensive missing features.
- This approach enhances the reliability of clinical decision support systems by using actual patient data.
- The method ensures transparency and interpretability in predicting clinical outcomes like suicide risk.
More Related Videos
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
12:18A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Related Concept Videos
Diagnostic and Statistical Manual of Mental Disorders (DSM)
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Introduction to Psychological Disorders
Deviant Behavior
Deviance in behavior refers to actions or thought patterns that significantly diverge from societal norms or...