Related Experiment Video
Updated: Jan 5, 2026

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Imputation and characterization of uncoded self-harm in major mental illness using machine learning
Praveen Kumar1,2, Anastasiya Nestsiarovich1, Stuart J Nelson3
1Center for Global Health, Department of Internal Medicine, University of New Mexico Health Sciences Center, Albuquerque, New Mexico, USA.
Machine learning effectively identifies unrecorded self-harm in individuals with major mental illness (MMI). This method reveals significant undercoding, particularly in males and older adults, highlighting risks for inadequate psychiatric care.
Area of Science:
- Psychiatry and Mental Health
- Health Informatics
- Machine Learning Applications
Background:
- Administrative claims data are crucial for understanding disease patterns but often underreport self-harm events in individuals with major mental illness (MMI).
- Accurate self-harm data are essential for effective mental healthcare planning and intervention.
Purpose of the Study:
- To impute uncoded self-harm events within administrative claims data for patients diagnosed with MMI.
- To characterize the incidence of self-harm and identify factors contributing to undercoding bias.
- To evaluate the performance of machine learning (ML) models in identifying self-harm events.
Main Methods:
- Utilized the IBM MarketScan database (2003-2016) encompassing over 10 million patients with MMI.
- Applied and compared five ML classifiers, selecting XGBoost for imputation on the full dataset.
- Validated ML model performance using data mislabeling techniques and comparison against a clinician-derived gold standard.
Main Results:
- ML imputation identified over 1.59 million self-harm events, vastly exceeding the 83,113 coded events, with high model accuracy (>0.99 AUC).
- The overall imputed self-harm incidence was 5.34%, significantly higher than the coded incidence of 0.28%.
- Self-harm undercoding was more prevalent in males and older individuals, and associated with specific diagnoses like substance abuse and mental health conditions.
Conclusions:
- Machine learning models can accurately recover underreported self-harm events in administrative claims data.
- Significant undercoding of self-harm exists for individuals with MMI, particularly affecting males and seniors.
- Improved data capture through ML can enhance psychiatric epidemiological research and inform targeted patient care.
More Related Videos
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
05:19Author Spotlight: Therapeutic Benefit of Closed-Loop Deep Brain Stimulation in Depression Treatment
Published on: July 7, 2023
Related Concept Videos
Self-Presentation: Self-Monitoring and Self-Handicapping
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Stereotype Content Model
Fundamental Attribution Error