Related Experiment Video
Updated: Jun 5, 2026

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
BUDS: Benchmark Uncertainty Design Selection for Two-Stage Single-Arm Phase II Trials
Rebecca Irlmeier1,2, Zhuoli Jin3,2, Fei Ye1,2
1Biostatistics and Bioinformatics Shared Resource, Sylvester Comprehensive Cancer Center, Miami, Florida.
Background:
Simon two-stage designs for binary endpoints and their time-to-event analogues, including the restricted-Kwak and Jung method, rely on a fixed historical benchmark. In practice, historical benchmarks are often uncertain due to small samples, population heterogeneity, changing eligibility criteria, and evolving standards of care. When the clinically relevant benchmark exceeds the value used to calibrate the design, Type I error rates can be substantially inflated, leading to costly advancement of ineffective treatments. Existing design-selection criteria generally optimize efficiency at a single benchmark without accounting for this uncertainty.
Methods:
We propose Benchmark Uncertainty Design Selection (BUDS), a framework for selecting two-stage designs for both binary and time-to-event (TTE) endpoints. Rather than relying on a single planning benchmark, BUDS addresses benchmark uncertainty by specifying a plausible range of response rates for binary endpoints or survival probabilities at a prespecified timepoint for TTE endpoints. Among feasible designs that satisfy both Type I error (α) and Type II error (β) constraints, BUDS-Worst-Regret minimizes the maximum difference in expected sample size from the benchmark-specific optimal design, whereas BUDS-Avg-EN minimizes average expected sample size across the range. By monotonicity, Type I error rate is controlled at the upper bound of the benchmark range.
Results:
Across representative scenarios, the Type I error rate increases substantially when true benchmarks exceed their planning values using single-benchmark objectives, reaching 0.407 for Simon Optimal design (response rates: 0.10 vs. 0.15; α = 0.10, β = 0.10) and 0.166 for restricted-Kwak and Jung design (survival probabilities: 0.50 vs. 0.67; α = 0.05, β = 0.20). In contrast, BUDS-selected designs maintain Type I error rate at or below the nominal level across the benchmark range, while the proposed objectives provide explicit criteria for prioritizing worst-case versus average efficiency.
Conclusions:
BUDS provides a transparent framework for selecting designs when historical benchmarks are uncertain. It makes the trade-off between robustness and efficiency explicit while controlling Type I error across the plausible benchmark range. The framework applies to both binary and time-to-event endpoints in two-stage single-arm Phase II trial planning and is implemented in the open-source BUDS R package and interactive Shiny app.
Related Concept Videos
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast, controlled...
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5% chance...
Propagation of Uncertainty from Random Error
Uncertainty: Overview
Uncertainty: Confidence Intervals
Testing a Claim about Standard Deviation
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
