Related Experiment Video
Updated: Jan 13, 2026

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
The importance of choosing a proper validation strategy in predictive models. Part 2: Recipes for (avoiding)
Eneko Lopez1, Giulia Gorla2, Jaione Etxebarria-Elezgarai3
1CIC nanoGUNE BRTA, Tolosa Hiribidea 76, San Sebastián, 20018, Spain; Department of Physics, University of the Basque Country (UPV/EHU), San Sebastián, 20018, Spain.
Overfitting in predictive modeling is often caused by poor validation and data issues, not just complexity. This study identifies common pitfalls and offers guidelines for trustworthy, generalizable models.
Area of Science:
- Machine Learning
- Predictive Modeling
- Data Science
Background:
- Overfitting is a significant challenge in predictive modeling, leading to poor generalization.
- It's often misattributed solely to model complexity, masking other critical issues.
Purpose of the Study:
- To identify overlooked causes of overfitting beyond model complexity.
- To provide practical guidelines for robust validation and trustworthy predictive models.
Main Methods:
- Analysis of common practices contributing to overfitting.
- Examination of data leakage and biased model selection.
- Review of publication pressures leading to overoptimization.
Main Results:
- Inadequate validation strategies are a primary driver of overfitting.
- Data preprocessing flaws and biased selection inflate apparent accuracy.
- Publication incentives can encourage result-driven overoptimization.
Conclusions:
- Addressing validation, preprocessing, and selection biases is crucial for reliable models.
- Researchers need practical guidelines to ensure model trustworthiness and generalizability.
- This work provides a blueprint for reproducible and robust predictive modeling.
Related Concept Videos
Data Validation
Key parameters for method validation include:
Data Validation
Nursing assessment guides are generally based on holistic models rather than medical...
Survival Tree
Building a Survival Tree
Constructing a...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
Improving Translational Accuracy

