Beyond Comparing Machine Learning and Logistic Regression in Clinical Prediction Modelling: Shifting from Model
Yanan Hu1, Xin Zhang2, Valerie Slavin3,4,5,6
1Monash Centre for Health Research and Implementation, Faculty of Medicine, Nursing and Health Sciences, Monash University, Melbourne, Australia.
Machine learning (ML) offers no universal advantage over logistic regression for clinical prediction models with tabular data. Improving data quality, not model complexity, is key to enhancing model reliability and real-world utility.
Area of Science:
- Clinical prediction modelling
- Machine learning in healthcare
- Statistical modelling
Background:
- Supervised machine learning (ML) is increasingly used for clinical prediction, especially for binary outcomes from tabular data.
- Debate exists on ML's advantage over traditional logistic regression in this domain.
- ML excels in unstructured data, but performance gains in structured clinical data are inconsistent.
Purpose of the Study:
- To synthesize comparative studies and simulation findings on ML versus logistic regression in clinical prediction.
- To argue against a universally superior modelling approach.
- To identify factors influencing model performance and suggest avenues for improvement.
Main Methods:
- Literature synthesis of recent comparative studies.
- Review of simulation findings.
- Viewpoint argument based on synthesized evidence.
Main Results:
- No single modelling approach (ML or logistic regression) is universally best for clinical prediction with tabular data.
- Model performance is highly dependent on dataset characteristics (e.g., linearity, sample size, predictor count, class proportion) and data quality (completeness, accuracy).
Conclusions:
- Efforts should focus on improving data quality rather than increasing model complexity.
- Enhanced data quality is more likely to improve the reliability and real-world utility of clinical prediction models.
More Related Videos
Related Concept Videos
Receiver Operating Characteristic Plot
Comparing the Survival Analysis of Two or More Groups
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.


