Related Experiment Video
Updated: Jun 26, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Dataset size versus homogeneity: A machine learning study on pooling intervention data in e-mental health dropout
Kirsten Zantvoort1, Nils Hentati Isacsson2, Burkhardt Funk1
1Institute of Information Systems, Leuphana University, Lueneburg, Germany.
Pooling data from internet-based cognitive behavioral therapy interventions can increase dataset sizes for machine learning. This approach improves prediction of intervention dropout for patients with depression, social anxiety, and panic disorder.
Area of Science:
- Digital mental health
- Machine learning in healthcare
- Psychological interventions
Background:
- Internet-based Cognitive Behavioral Therapy (iCBT) interventions often face challenges with limited dataset sizes.
- Small datasets can hinder the development and accuracy of machine learning models for predicting patient outcomes.
- Understanding user behavior and symptom data similarities across different iCBT interventions is crucial for data aggregation.
Purpose of the Study:
- To investigate the feasibility and benefits of pooling data from distinct iCBT interventions.
- To examine user behavior and symptom data similarities among iCBT interventions for depression, social anxiety, and panic disorder.
- To determine if pooled data enhances the accuracy of predicting intervention dropout.
Main Methods:
- Analysis of routine care patient data (n=6418) from Internet Psychiatry in Stockholm.
- Application of clustering techniques to identify patient groups based on activity levels.
- Development and comparison of dropout prediction models using individual versus pooled intervention datasets, tested on varying dataset sizes.
Main Results:
- Clustering revealed three distinct patient groups characterized by activity levels, independent of specific interventions.
- Pooling intervention data significantly improved dropout prediction accuracy in eight out of nine tested scenarios.
- Models trained on smaller datasets demonstrated a tendency to overestimate prediction results.
Conclusions:
- Patients with depression, social anxiety, and panic disorder exhibit similar online activity and dropout patterns across iCBT interventions.
- Pooling data from different iCBT interventions is a viable strategy to overcome small dataset limitations in psychological research.
- This approach can enhance the robustness and generalizability of machine learning models in digital mental health.
More Related Videos
Related Concept Videos
Regression Toward the Mean
Longitudinal Research
Longitudinal Studies
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Mechanistic Models: Compartment Models in Individual and Population Analysis

