Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Sample Size Calculation01:19

Sample Size Calculation

3.3K
Knowledge of the sample size is the first requirement to conduct random sampling or an experiment. The sample size is the total number of units, observations, or groups (in some cases) used to get the data to estimate a population parameter. As the name suggests, the sample size is that of the sample drawn from the population and differs from the population size.
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
3.3K
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches01:23

Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches

125
Biopharmaceutical studies constitute a vital field aiming to enhance drug delivery methods and refine therapeutic approaches, drawing upon diverse interdisciplinary knowledge. In research methodologies, the choice between controlled and non-controlled studies significantly influences the study's reliability and accuracy.
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
125
Survival Tree01:19

Survival Tree

79
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
 Building a Survival Tree
Constructing a...
79
Odds Ratio01:09

Odds Ratio

123
The odds ratio (OR) is a statistical measure used extensively in epidemiology and research to quantify the strength of association between exposure and outcome across different groups. Unlike relative risk, which compares the probabilities of an event occurring, the odds ratio compares the odds of an event occurring in the exposed group to the odds of it occurring in the unexposed group. The odds, in this context, are calculated as the probability of the event happening divided by the...
123
Statistical Methods for Analyzing Epidemiological Data01:25

Statistical Methods for Analyzing Epidemiological Data

347
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
347
Comparing the Survival Analysis of Two or More Groups01:20

Comparing the Survival Analysis of Two or More Groups

175
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
175

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Mode-Selective Dual-Level Vibrational Perturbation Theory Assisted by Machine Learning for Rotational and Vibrational Spectra of Benzoic Acid and Aspirin.

The journal of physical chemistry. A·2026
Same author

Assessment of Adverse Events Using the Therapy-Disability-Neurology (TDN) Grading System in a Cohort of Aneurysmal Subarachnoid Hemorrhage Patients: A Single-Center Retrospective Cohort Study.

Brain sciences·2026
Same author

Reduced Indocyanine Green Clearance Is Associated with Enteral Feeding Intolerance in Septic Patients Without Overt Liver Injury.

Journal of clinical medicine·2026
Same author

IL-22BP attenuates right ventricular remodeling in pulmonary arterial hypertension.

Clinical science (London, England : 1979)·2026
Same author

Patient-reported non-motor outcomes after endovascular thrombectomy and intravenous thrombolysis: an observational study.

European stroke journal·2026
Same author

Oxytocin modulates the neurocomputational mechanisms engaged in learning rank relationships in social networks.

Proceedings of the National Academy of Sciences of the United States of America·2026

Related Experiment Video

Updated: Jun 21, 2025

An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.0K

An evaluation of sample size requirements for developing risk prediction models with binary outcomes.

Menelaos Pavlou1, Gareth Ambler2, Chen Qu2

  • 1Department of Statistical Science, UCL, London, UK. m.pavlou@ucl.ac.uk.

BMC Medical Research Methodology
|July 10, 2024
PubMed
Summary

Existing formulas for sample size calculation in risk prediction models are unreliable for high model strengths. A new simulation-based approach is proposed to accurately estimate sample sizes for improved clinical decision-making.

Keywords:
CalibrationDiscriminationSample sizeSimulation

More Related Videos

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.5K
Establishing a Competing Risk Regression Nomogram Model for Survival Data
04:57

Establishing a Competing Risk Regression Nomogram Model for Survival Data

Published on: October 23, 2020

10.1K

Related Experiment Videos

Last Updated: Jun 21, 2025

An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.0K
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.5K
Establishing a Competing Risk Regression Nomogram Model for Survival Data
04:57

Establishing a Competing Risk Regression Nomogram Model for Survival Data

Published on: October 23, 2020

10.1K

Area of Science:

  • Biostatistics
  • Clinical Epidemiology
  • Health Informatics

Background:

  • Risk prediction models are crucial for clinical decision-making.
  • Small sample sizes can compromise model performance and generalizability.
  • Calibration slope (CS) and mean absolute prediction error (MAPE) are key metrics for sample size calculations.

Purpose of the Study:

  • To evaluate the performance of existing sample size calculation formulas for risk prediction models.
  • To assess the accuracy of these formulas under various conditions, including different model strengths and outcome prevalences.
  • To propose an improved method for sample size estimation in clinical risk prediction.

Main Methods:

  • A simulation study was conducted to evaluate two proposed sample size calculation formulas.
  • The study analyzed the formulas' performance based on anticipated data features like outcome prevalence and c-statistic.
  • The evaluation considered both binary outcomes and time-to-event data with censoring.

Main Results:

  • Existing formulas perform adequately for models with lower strength (c-statistic < 0.8).
  • The CS formula underestimates required sample size for high model strengths (c-statistic > 0.8), necessitating increases of 50-100%.
  • The MAPE formula tends to overestimate sample size for high model strengths, with effects more pronounced at higher outcome prevalences.

Conclusions:

  • Current sample size formulas are generally appropriate for lower model strengths but biased for higher strengths common in clinical settings.
  • A novel simulation-based approach, implemented in the R package 'samplesizedev', is proposed for accurate sample size estimation.
  • The proposed method accounts for model stability by calculating variability in CS and MAPE.