Related Experiment Video
Updated: Apr 28, 2026

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Cross-study validation for the assessment of prediction algorithms
Christoph Bernau1, Markus Riester1, Anne-Laure Boulesteix2
1Leibniz Supercomputing Center, Garching, Department for Medical Informatics, Biometry and Epidemiology, Munich, Germany, Cambridge, MA, Dana-Farber Cancer Institute, Boston, Harvard School of Public Health, Boston, USA and City University of New York School of Public Health, Hunter College, New York, USALeibniz Supercomputing Center, Garching, Department for Medical Informatics, Biometry and Epidemiology, Munich, Germany, Cambridge, MA, Dana-Farber Cancer Institute, Boston, Harvard School of Public Health, Boston, USA and City University of New York School of Public Health, Hunter College, New York, USA.
Cross-study validation offers a more realistic performance evaluation for high-dimensional prediction models than standard cross-validation. This approach is crucial for ensuring accurate predictions in real-world applications, especially in complex datasets like those found in cancer research.
Area of Science:
- Biostatistics
- Machine Learning
- Bioinformatics
Background:
- High-dimensional prediction models are prevalent in statistical and machine-learning literature.
- Conventional cross-validation on exemplary datasets may not accurately reflect real-world prediction performance.
- Accurate prediction for independent samples in diverse settings is the primary goal in many applications.
Purpose of the Study:
- To introduce and implement a systematic 'cross-study validation' approach.
- To evaluate high-dimensional prediction models using independent datasets.
- To compare cross-study validation with conventional cross-validation.
Main Methods:
- Developed and implemented a cross-study validation framework.
- Utilized simulations and a collection of eight breast cancer gene-expression datasets.
- Computed C-index for pairwise training and validation datasets to assess distant metastasis-free survival (DMFS).
Main Results:
- Standard cross-validation yielded inflated discrimination accuracy compared to cross-study validation across all tested algorithms.
- The ranking of learning algorithms varied significantly between cross-validation and cross-study validation.
- Algorithms optimal under cross-validation may be suboptimal for independent validation.
Conclusions:
- Cross-study validation provides a more reliable assessment of prediction model performance in high-dimensional settings.
- The findings highlight the limitations of conventional cross-validation for real-world applicability.
- The proposed method is essential for selecting robust prediction algorithms for independent data.
Related Concept Videos
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Receiver Operating Characteristic Plot
Reliability and Validity
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Improving Translational Accuracy

