Related Experiment Video
Updated: Jan 13, 2026

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
The importance of choosing a proper validation strategy in predictive models. Part 2: Recipes for (avoiding)
Eneko Lopez1, Giulia Gorla2, Jaione Etxebarria-Elezgarai3
1CIC nanoGUNE BRTA, Tolosa Hiribidea 76, San Sebastián, 20018, Spain; Department of Physics, University of the Basque Country (UPV/EHU), San Sebastián, 20018, Spain.
Abstract:
Overfitting remains one of the most pervasive and deceptive pitfalls in predictive modeling. It leads to models that perform exceptionally well on training data but cannot be transferred nor generalized to real-world scenarios. Although overfitting is usually attributed to excessive model complexity, it is often the result of inadequate validation strategies, faulty data preprocessing and biased model selection, problems that can inflate apparent accuracy and compromise predictive reliability. In this second part of our series, we examine the most common yet overlooked practices that contribute to overfitting, ranging from data leakage in preprocessing to the pressures of scientific publishing that encourage result-driven overoptimization. By identifying these pitfalls and providing practical guidelines for performing robust validation protocols, this work serves as a blueprint for researchers to ensure their models are not only high-performing but also trustworthy, reproducible, and generalizable.
Related Concept Videos
Data Validation
Key parameters for method validation include:
Data Validation
Nursing assessment guides are generally based on holistic models rather than medical...
Survival Tree
Building a Survival Tree
Constructing a...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
Improving Translational Accuracy

