Related Experiment Video
Updated: May 30, 2025

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Extensive benchmarking of a method that estimates external model performance from limited statistical
Tal El-Hay1, Jenna M Reps2, Chen Yanover3
1KI Research Institute, Kfar Malal, Israel. talelh@kinstitute.org.il.
Estimating external model performance using only summary statistics accurately predicts how models will perform on new data. This method facilitates the deployment of predictive models by enabling external validation without patient-level data.
Area of Science:
- Medical Informatics
- Health Services Research
- Biostatistics
Background:
- External validation is crucial for assessing predictive model generalizability.
- Access to external patient-level data for validation is often restricted.
- Existing methods for external validation can be data-intensive.
Purpose of the Study:
- To benchmark a novel method for estimating external model performance using only summary statistics.
- To evaluate the accuracy of this method across various metrics and datasets.
- To demonstrate the feasibility of assessing model transportability with limited external data.
Main Methods:
- The proposed method was benchmarked using five large, heterogeneous US data sources.
- Each data source served sequentially as the internal training set, with others acting as external validation sets.
- Performance was estimated using only external summary statistics.
Main Results:
- The method accurately estimated external model performance across multiple metrics.
- 95th percentile errors for discrimination, calibration, and overall accuracy were low (0.03, 0.08, 0.0002, and 0.07).
- These findings confirm the feasibility of estimating model transportability.
Conclusions:
- Estimating external model performance using summary statistics is feasible and accurate.
- This approach can accelerate the deployment of predictive models by simplifying external validation.
- The method offers a practical solution for assessing model generalizability with limited data access.
More Related Videos
04:35Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach
Published on: July 3, 2020
20:24Characterization of Complex Systems Using the Design of Experiments Approach: Transient Protein Expression in Tobacco as a Case Study
Published on: January 31, 2014
Related Concept Videos
Mechanistic Models: Compartment Models in Individual and Population Analysis
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Testing a Claim about Standard Deviation
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
What are Estimates?
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such...
Quantifying and Rejecting Outliers: The Grubbs Test
Statistical Analysis: Overview
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...