Related Experiment Videos
Adaptive Gaussian process search for simulation-based sample size estimation in clinical prediction models:
Oyebayo Ridwan Olaniran1,2, Diana Shamsutdinova3,4, Sarah Markham3
1Department of Biostatistics & Health Informatics, King's College London, London, UK. ridwan.olaniran@kcl.ac.uk.
Background:
Determining an adequate sample size is essential for developing reliable and generalisable clinical prediction models, yet practical guidance on selecting appropriate sample size calculation methods remains limited. Existing analytical and simulation-based tools impose restrictive assumptions and focus on mean-based criteria. We present and validate pmsims, an R package that uses Gaussian process (GP) surrogate modelling to provide a flexible and efficient simulation-based framework for sample size determination, validated here across binary, continuous, and survival prediction modelling contexts.
Methods:
We conducted a comprehensive simulation study with two aims. Aim 1 compared three search engines implemented in pmsims, a GP surrogate-based adaptive procedure (gp), a deterministic bisection method (bisection), and a hybrid GP-bisection approach (gp-bs), across binary, continuous, and survival outcomes. Scenarios varied outcome prevalence or event rate, predictor dimensionality ([Formula: see text]), target performance metric (discrimination and calibration slope), aggregation criterion (mean vs 80% assurance), and total simulation budget ([Formula: see text]). Each scenario was replicated 100 times; estimator stability was assessed via the coefficient of variation (CV). Aim 2 benchmarked the best-performing pmsims engine against pmsampsize (analytical) and samplesizedev (simulation-based) across a wider range of realistic scenarios, evaluating recommended sample sizes, computational time, and achieved model performance on independent validation datasets of 30,000 observations.
Results:
The GP-based search engine consistently yielded the most stable sample size estimates (lowest CV) across all outcome types, ranking highest in 9/12 outcome-aggregation metric configurations. Its advantage was most pronounced in low-signal, high-dimensional settings, and was accentuated with [Formula: see text] replications per evaluation and a budget of [Formula: see text], after which gains were minimal. In benchmark comparisons, pmsims (mean) achieved performance deviations within [Formula: see text] of the prespecified target across binary, continuous, and survival outcomes, comparable to samplesizedev and substantially outperforming pmsampsize in high-discrimination settings (deviations up to [Formula: see text]).
Conclusions:
The pmsims package, using the GP-based search engine with [Formula: see text] replications per evaluation and a budget of [Formula: see text], provides a computationally efficient and flexible framework for principled sample size planning in clinical prediction modelling. It reliably achieves performance targets across the range of standard-model scenarios evaluated while requiring fewer model evaluations than non-adaptive simulation approaches, offering a compelling alternative to both analytical formulae and exhaustive simulation-based search.
Related Concept Videos
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Statistical Software for Data Analysis and Clinical Trials
Parametric Survival Analysis: Weibull and Exponential Methods
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Estimating Population Standard Deviation
Distributions to Estimate Population Parameter
Mechanistic Models: Compartment Models in Individual and Population Analysis