Related Experiment Video
Updated: Jun 5, 2026

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Integrating epidemiologic modeling and explainable machine learning to predict and identify factors associated with
Mustapha Aliyu Muhammad1, Jamilu Sani2, Salad Halane3
1Department of Biostatistics and Epidemiology, College of Public Health, East Tennessee State University, Johnson City, TN, USA.
Background:
Depression is a major public health concern, with Tennessee ranking among the U.S. states with the highest prevalence. Despite its burden, many cases remain undetected due to limited screening and access to mental health services. This study integrated epidemiologic modelling and explainable ML techniques to predict self-reported depression and identify key risk factors among Tennessee adults using the Behavioral Risk Factor Surveillance System (BRFSS) 2023 data.
Methods:
We conducted a cross-sectional analysis of 5596 adults from the 2023 Tennessee BRFSS, representing 5,569,707 weighted respondents. The primary outcome was lifetime self-reported diagnosis of depression. The Oregon BRFSS 2023 was used as an external validation dataset. Eight machine learning algorithms were trained using 5-fold stratified cross-validation. Model performance was evaluated using AUROC, PR-AUC, accuracy, precision, recall, F1-score, balanced accuracy, DeLong's test, and McNemar's test, while model interpretability was assessed using SHapley Additive exPlanations (SHAP).
Results:
The weighted prevalence of self-reported depression among Tennessee adults was 27.3%. Among the evaluated algorithms, XGBoost, Gradient Boosting, Random Forest, and Logistic Regression demonstrated the strongest and highly comparable external validation performance. DeLong's test for AUROC and paired bootstrap resampling for PR-AUC showed no statistically significant differences among these four leading models. McNemar's test produced a similar pattern for paired classification errors. SHAP interpretation identified sex, ACEs category, memory decline, disability category, race/ethnicity, poor physical activity, and age group as the most influential predictors of self-reported depression.
Conclusions:
This study demonstrates the utility of integrating explainable machine learning approaches to predict and identify factors associated with self-reported depression, thereby enhancing the use of public health surveillance systems in early identification of high-risk populations and informing targeted mental health interventions.
Related Concept Videos
Depressive Disorders: Etiology
Biological Factors in Depression
Biological predispositions significantly influence the risk of developing depressive disorders. Genetic studies highlight the role of variations in the serotonin transporter...
Steps in Outbreak Investigation
Mechanistic Models: Compartment Models in Individual and Population Analysis
Statistical Methods for Analyzing Epidemiological Data
Models of Health Promotion and Illness Prevention I
The health belief model (HBM) attempts to predict health-related behavior in specific belief patterns. According to the HBM, a person's...
Depression: Overview