Related Experiment Video
Updated: Jul 8, 2025

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
Building a better model: abandon kitchen sink regression
Stefan Kuhle1,2, Mary Margaret Brown3, Sanja Stanojevic4
1Institute for Medical Biostatistics, Epidemiology and Informatics (IMBEI), Johannes Gutenberg University Mainz, Mainz, Germany stefan.kuhle@uni-mainz.de.
Abstract:
This paper critically examines 'kitchen sink regression', a practice characterised by the manual or automated selection of variables for a multivariable regression model based on p values or model-based information criteria. We highlight the pitfalls of this method, using examples from perinatal/neonatal medicine, and propose more robust alternatives. The concept of directed acyclic graphs (DAGs) is introduced as a tool for describing and analysing causal relationships. We highlight five key issues with 'kitchen sink regression': (1) the disregard for the directionality of variable relationships, (2) the lack of a meaningful causal interpretation of effect estimates from these models, (3) the inflated alpha error rate due to multiple testing, (4) the risk of overfitting and model instability and (5) the disregard for content expertise in model building. We advocate for the use of DAGs to guide variable selection for models that aim to examine associations between a putative risk factor and an outcome and emphasise the need for a more thoughtful and informed use of regression models in medical research.
Related Concept Videos
Improving Translational Accuracy
Regression Toward the Mean
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Quantifying and Rejecting Outliers: The Grubbs Test
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...

