Related Experiment Videos
Power analysis in randomized clinical trials based on item response theory
Rebecca Holman1, Cees A W Glas, Rob J de Haan
1Department of Clinical Epidemiology and Biostatistics, Academic Medical Center, Amsterdam, The Netherlands. R.Holman@amc.uva.nl
Controlled Clinical Trials
|July 17, 2003
Summary
Designing randomized clinical trials (RCTs) requires careful consideration of sample size and instrument length. Using more items in patient-reported outcome questionnaires generally reduces the number of participants needed for detecting treatment effects, especially small ones.
Area of Science:
- Psychometrics
- Clinical Trials Methodology
- Biostatistics
Background:
- Patient-reported outcomes (PROs) are increasingly vital endpoints in randomized clinical trials (RCTs).
- Item Response Theory (IRT) is gaining traction for analyzing questionnaire data in clinical research.
- Understanding the interplay between sample size, instrument length, and statistical power in RCTs is crucial for efficient trial design.
Purpose of the Study:
- To investigate the behavior of a test statistic for comparing latent trait levels between two groups using the two-parameter logistic IRT model.
- To examine how the number of patients per arm, number of items, and statistical power interact in RCTs.
- To provide guidance on instrument selection for detecting various treatment effect sizes in IRT-analyzed RCTs.
Main Methods:
- A simulation study was conducted to evaluate the performance of an IRT-based test statistic.
- The simulation examined the two-parameter logistic IRT model for binary data.
- The study analyzed the relationship between sample size, number of items, and power to detect minimal, moderate, and substantial effects.
Main Results:
- The number of patients required in an RCT arm is influenced by the number of items used.
- With at least 20 items, the number of items has minimal impact on detecting moderate (0.5) to substantial (0.8) effects with 80% power.
- Detecting small effects (0.2) is more sensitive to the number of items; using 5 items requires 950 patients/arm, while 50 items require 450 patients/arm for 80% power.
- Analysis of SF-36, SF-12, and SF-8 instruments yielded slightly different results due to multi-category items.
Conclusions:
- For RCTs aiming to detect small effects using IRT, utilizing very short instruments is not advisable.
- Instrument length significantly impacts the required sample size for detecting small treatment effects.
- Trial designers should carefully select the number of items based on the expected effect size and desired statistical power.