Interpreting Effect Size of Patient-Reported Outcome Measures in Epilepsy: Towards Empirical Benchmarks
Ann Subota1, Mandavi Kashyap2, Samuel Wiebe3
1Department of Clinical Neurosciences, Cumming School of Medicine, University of Calgary, 3330 Hospital Drive NW, Calgary, AB, Canada, T2N 4N1.
Objective:
To enhance effect size (ES) interpretation in epilepsy patient-reported outcome measures, focusing on single-item global rating (SIGR) and multi-item scales (MIS), by deriving evidence-based thresholds for small, medium, and large ES, and comparing them to conventional Cohen's ES thresholds.
Methods:
Data were collected from articles identified from a prior scoping review, which followed Preferred Reporting Items for Systematic reviews and Meta-Analyses (PRISMA) and Joanna Briggs Institute guidelines. CINAHL, Embase, MEDLINE, PsycINFO, and the Cochrane Register of Controlled Trials databases were searched. English-language articles with ≥30 persons with epilepsy that used at least one SIGR and one MIS were included. Correlation coefficients were transformed to Fisher's z and back transformed to r. Cohen's d values were converted to Hedges' g using standard formulas. Thresholds were analysed separately for correlations (r), and for standardized mean differences (Cohen's d). For each metric, distribution of absolute ES was characterized, and 25th, 50th, 75th percentiles were estimated using nonparametric bootstrap resampling at the study level. These percentiles defined "small", "medium" and "large" ES, respectively. Context-specific empirical ES thresholds were also derived.
Results:
For small and medium effects, omnibus empirical thresholds for correlations (r) (0.21, 0.35) are larger than conventional cut-offs (0.10, 0.30). Compared to Cohen's d thresholds, the omnibus empirical small ES was larger (0.24 vs 0.20), and the medium and large ES were smaller (0.44 vs 0.50, 0.78 vs 0.80). Empirical thresholds for d were generally smaller than those for r. Empirical thresholds for both r and d varied widely depending on specific contexts.
Significance:
Conventional Cohen's benchmarks for small, medium and large ES for r and d rarely align with empirically derived thresholds for PROMs in epilepsy. A single set of empirical ES thresholds is simplistic. Wide variation in context-specific ES thresholds should be considered in study design and interpretation of study results, as they can directly impact statistical and clinical significance. This work outlines methodological frameworks, comparative insights between SIGRs and MIS, and demonstrates context-specific ES thresholds provide a more informative framework that applying a single universal set of benchmarks, which can be directly applied to other therapeutic areas where similar data exists in future studies.


