Related Experiment Videos
Depression prediction and key factors: A comparative analysis of logistic regression and machine learning models
Kripa Josten1, Vennila Jaganathan2
1Research Scholar, Manipal College of Health Professions, Manipal Academy of Higher Education, Manipal, Karnataka, India.
Objectives:
Factors associated with depression were explored in this study through logistic regression, and predictive performance was compared with various Machine Learning models.
Methods:
WHO SAGE India wave 2 data were used with depression as the outcome variable. and predictors were sociodemographic, health, and psychosocial variables. Descriptive analysis and Logistic regression were estimated. Random Forest, XGBoost, Support Vector Machine, Logistic Regression, Bagging, Decision Tree, Naïve Bayes, Ridge Logistic Regression, Neural Networks, and K Nearest Neighbors are the ten Machine Learning algorithms that were used. Performance measures consisted of accuracy, Area Under Curve, precision, recall, F1 score, Hamming loss, Jaccard score, and Matthew's correlation coefficient. Random Forest and XGBoost were used to assess feature importance.
Results:
Depression was also more prevalent among younger adults, women, and individuals with poor self-rated health, stress, and sleep disturbances. Logistic regression revealed age and feeling low or sad as a factor (p = 0.008, p = 0.021). Most models demonstrated only moderate discriminative ability, with the AUC below 0.70, with better-performing models being Ridge regression (AUC = 0.716) and Random Forest (AUC = 0.713). Feature importance universally identified age, perception of health, quality of life, and depressive symptoms as important predictors.
Conclusions:
Logistic regression provides interpretability, and Machine Learning increases predictive accuracy. Combining both can enhance depression prediction and screening in public health practice.
Related Concept Videos
Depressive Disorders: Etiology
Biological Factors in Depression
Biological predispositions significantly influence the risk of developing depressive disorders. Genetic studies highlight the role of variations in the serotonin transporter...
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...