Related Experiment Video
Updated: Jun 5, 2026

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
Widespread use of invalid statistical tests in biomedical machine learning.
Tianchu Zeng1,2,3,4,5,6,7, Hetu Li1,3,4,5,6,7, Shaoshi Zhang1,2,3,4,5,6,7,8
1Centre for Sleep & Cognition & Centre for Translational Magnetic Resonance Research, Yong Loo Lin School of Medicine, National University of Singapore, Singapore.
Most biomedical machine learning studies incorrectly compare model performance by ignoring cross-validation fold dependence, inflating false positives. A new SHARP test offers valid comparisons, improving reliability in research.
Area of Science:
- Biomedical research
- Machine learning
- Statistical inference
Background:
- Machine learning accelerates biomedical research, with cross-validation commonly used for performance comparison.
- Standard statistical tests for comparing prediction performance assume independence, which is violated by cross-validation folds.
- Ignoring this fold dependence inflates false positive rates, compromising research validity.
Purpose of the Study:
- To investigate the prevalence and impact of ignoring cross-validation fold dependence in biomedical machine learning.
- To propose a novel statistical method for valid model comparison that accounts for fold dependence.
Main Methods:
- A PRISMA-guided meta-analysis of 210 studies was conducted to assess reporting practices.
- Simulations across 420 scenarios with diverse datasets evaluated the impact of fold dependence and tested new methods.
- The proposed SHARP (Split-HAlf RePeated) test was developed as a modification to standard cross-validation.
Main Results:
- 97% of studies ignored fold dependence when comparing prediction performance, a problem pervasive across fields.
- Ignoring fold dependence leads to invalid false positive control, especially with repeated cross-validation.
- The SHARP test demonstrated superior performance in controlling false positives, maintaining statistical power, and calibrating confidence intervals compared to 12 other tests.
Conclusions:
- The widespread failure to account for cross-validation fold dependence undermines the reliability of machine learning research in biomedicine.
- The SHARP test offers a robust solution for valid statistical inference in model comparison.
- Best practices and reporting guidelines are needed to ensure accurate model comparison in biomedical machine learning.
Related Concept Videos
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance, comparing...
Statistical Software for Data Analysis and Clinical Trials
Regression Toward the Mean
Statistical Significance
Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with data...
Errors In Hypothesis Tests