Related Experiment Video
Updated: Jun 15, 2026

Author Spotlight: Therapeutic Benefit of Closed-Loop Deep Brain Stimulation in Depression Treatment
Published on: July 7, 2023
Multi-site benchmark classification of major depressive disorder using machine learning on cortical and subcortical
Vladimir Belov1, Tracy Erwin-Grabner1, Moji Aghajani2,3
1Laboratory of Systems Neuroscience and Imaging in Psychiatry (SNIP-Lab), Department of Psychiatry and Psychotherapy, University Medical Center Göttingen (UMG), Georg-August University, Von-Siebold-Str. 5, 37075, Göttingen, Germany.
Machine learning models achieved ~62% accuracy in classifying major depressive disorder (MDD) using neuroimaging data from over 5,000 individuals. Harmonizing data reduced accuracy to ~52%, highlighting challenges in generalizable MDD classification.
Area of Science:
- Neuroimaging
- Computational Psychiatry
- Machine Learning
Background:
- Machine learning (ML) shows promise for classifying neuropsychiatric disorders using neuroimaging data.
- Existing ML algorithms face limitations such as small sample sizes, data leakage, and overfitting, hindering diagnostic predictive power.
Purpose of the Study:
- To establish a generalizable machine learning classification benchmark for major depressive disorder (MDD).
- To overcome limitations of previous studies by utilizing the largest multi-site sample size to date (N=5365).
- To evaluate shallow linear and non-linear models for MDD classification.
Main Methods:
- Utilized brain measures from standardized ENIGMA analysis pipelines in FreeSurfer.
- Employed shallow linear and non-linear machine learning models.
- Classified major depressive disorder (MDD) versus healthy controls (HC) using a large, multi-site dataset (N=5365).
- Assessed the impact of data harmonization techniques (e.g., ComBat) on classification accuracy.
Main Results:
- Achieved a balanced accuracy of approximately 62% for classifying MDD versus HC before data harmonization.
- Balanced accuracy dropped to approximately 52% after harmonizing the data using ComBat.
- Classification accuracy approached random chance levels in stratified subgroups based on clinical variables (age of onset, medication, episode count, sex).
Conclusions:
- Current shallow machine learning models using standard neuroimaging features provide limited generalizable classification accuracy for MDD, even with large sample sizes.
- Data harmonization techniques may reduce classification performance, indicating potential loss of relevant information or introduction of biases.
- Future research should explore higher-dimensional features and advanced machine/deep learning methods for improved MDD classification.

