Related Experiment Video
Updated: Nov 9, 2025

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Data-driven methods distort optimal cutoffs and accuracy estimates of depression screening tools: a simulation study
Parash Mani Bhandari1, Brooke Levis2, Dipika Neupane1
1Lady Davis Institute for Medical Research, Jewish General Hospital, Montreal, Quebec, Canada; Department of Epidemiology, Biostatistics and Occupational Health, McGill University, Montreal, Quebec, Canada.
Data-driven methods in small accuracy studies can lead to incorrect optimal cutoffs and biased accuracy estimates. Larger sample sizes improve the reliability of these estimates for screening tools like the Edinburgh Postnatal Depression Scale (EPDS).
Area of Science:
- Psychometrics
- Biostatistics
- Clinical Trial Design
Background:
- Accurate screening tools are crucial for timely intervention in conditions like postpartum depression.
- The Edinburgh Postnatal Depression Scale (EPDS) is widely used, but its optimal cutoff may vary with sample characteristics.
- Data-driven methods are often employed to determine optimal cutoffs and estimate accuracy in real-world studies.
Purpose of the Study:
- To assess the impact of sample size on the accuracy of data-driven methods for determining optimal cutoffs.
- To evaluate the bias in accuracy estimates (sensitivity and specificity) when using smaller sample sizes compared to population-level data.
- To compare optimal cutoffs derived from simulated samples with the established population optimal cutoff for the EPDS.
Main Methods:
- Simulated 1,000 samples each for sizes 100, 200, 500, and 1,000 from a large database (n=13,255) of EPDS scores.
- Determined optimal cutoffs by maximizing Youden's J statistic (sensitivity + specificity - 1) in each simulated sample.
- Compared optimal cutoffs and accuracy estimates from simulated samples against population values.
Main Results:
- Optimal cutoffs varied significantly across sample sizes, ranging from ≥5 to ≥17 for n=100 and ≥8 to ≥13 for n=1,000.
- The percentage of samples identifying the population optimal cutoff (≥11) increased with sample size: 30% (n=100) to 71% (n=1,000).
- Mean overestimation of sensitivity was highest in smaller samples (6.5 pp for n=100) and decreased with larger samples (1.4 pp for n=1,000). Specificity was generally underestimated.
Conclusions:
- Data-driven methods in small accuracy studies can yield inaccurate optimal cutoffs for screening tools like the EPDS.
- Accuracy estimates, particularly sensitivity, can be significantly overstated in smaller sample sizes.
- Larger sample sizes are essential for reliable determination of optimal cutoffs and accurate estimation of screening tool performance.
More Related Videos
05:19Author Spotlight: Therapeutic Benefit of Closed-Loop Deep Brain Stimulation in Depression Treatment
Published on: July 7, 2023
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Related Concept Videos
Censoring Survival Data
Regression Toward the Mean
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Blind Procedures
Bias in Epidemiological Studies