Related Experiment Videos
Interpretable machine learning for depression symptom classification in NHANES: Performance in a curated
Omid Salimi1,2, Zahra Ghahramani3, Mahsa Taghavi Zenouz4
1Department of Medicine, Islamic Azad University, Najafabad, Iran.
Background:
Depressive disorders are among the most common psychiatric conditions worldwide and frequently remain undetected or misinterpreted. Scalable computational tools may help characterize depressive-symptom patterns in large health datasets, but their clinical use requires realistic validation and calibration.
Objective:
We evaluated machine-learning models for classifying PHQ-9 depressive-symptom status in NHANES and compared performance in a curated high-confidence corpus and the full analytic population.
Methods:
This cross-sectional analysis included 37,959 NHANES participants with valid PHQ-9 data and prespecified demographic, sleep, and dietary predictors. The primary outcome was mild-or-greater depressive symptoms, defined as PHQ-9 ≥ 5. A balanced high-confidence corpus was constructed using clear PHQ-9 rules, teacher-model confidence ranking, Isolation Forest filtering, and class-specific resampling. A LightGBM classifier was compared with logistic regression, random forest, gradient boosting, support vector machine, XGBoost, k-nearest neighbors, and Gaussian naïve Bayes. Performance, calibration, temporal validation, sensitivity analyses, and SHAP-based interpretability were assessed.
Results:
The curated corpus contained 5000 balanced observations. LightGBM achieved near-perfect curated-holdout performance, but full-population performance was moderate: accuracy 0.678, F1-score 0.503, precision 0.407, recall 0.659, and ROC-AUC 0.726. Comparator models showed similar full-population discrimination. Removing sleep predictors reduced ROC-AUC to 0.611. Full-population calibration was poor, and self-reported trouble sleeping was the dominant contributor.
Conclusion:
Interpretable machine learning identified reproducible depressive-symptom patterns in NHANES, largely driven by sleep variables. However, modest precision and poor calibration preclude stand-alone clinical use without external validation and recalibration.