Related Experiment Video
Updated: Oct 16, 2025

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Risk of bias in studies on prediction models developed using supervised machine learning techniques: systematic
Constanza L Andaur Navarro1,2, Johanna A A Damen3,2, Toshihiko Takada3
1Julius Centre for Health Sciences and Primary Care, University Medical Center Utrecht, Utrecht University, Utrecht, Netherlands c.l.andaurnavarro@umcutrecht.nl.
Most machine learning prediction models in medicine have poor quality and high risk of bias due to small study size and inadequate data handling. Improving study design and validation is crucial for clinical application.
Area of Science:
- Medical Informatics
- Machine Learning in Healthcare
- Clinical Prediction Models
Background:
- Machine learning (ML) is increasingly used for developing diagnostic and prognostic prediction models in medicine.
- Assessing the methodological quality and risk of bias in these studies is essential for reliable clinical application.
Purpose of the Study:
- To systematically review and assess the methodological quality and risk of bias of studies developing prediction models using machine learning techniques across all medical specialties.
Main Methods:
- A systematic review of studies published between January 2018 and December 2019 was conducted.
- The Prediction Risk Of Bias Assessment Tool (PROBAST) was used to evaluate methodological quality and risk of bias in 152 developed models and 19 external validations.
- Risk of bias was assessed across domains: participants, predictors, outcome, and analysis.
Main Results:
- Out of 171 analyses, 148 (87%) were rated at high risk of bias, with the analysis domain most frequently affected.
- Key issues included inadequate events per predictor (56%), poor handling of missing data (41%), and improper assessment of overfitting (39%).
- While data sources were generally appropriate, information on blinding was often absent.
Conclusions:
- The majority of studies developing ML-based prediction models exhibit poor methodological quality and a high risk of bias.
- Addressing issues like small study size, missing data, and overfitting is critical.
- Improving the design, conduct, reporting, and validation of these studies is necessary to facilitate their integration into clinical practice.
More Related Videos
Related Concept Videos
Bias
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
Bias in Epidemiological Studies
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Survival Tree
Building a Survival Tree
Constructing a...
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Regression Toward the Mean

