Related Experiment Video
Updated: May 2, 2026

07:35
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
7.6K
A supervised machine learning approach with feature selection for sex-specific biomarker prediction.
Luke Meyer1, Danielle Mulder2, Joshua Wallace1
17 Long Tom Place Kanonberg Bellville Western Cape, Siriuz Pty Ltd., Cape Town, South Africa.
NPJ Systems Biology and Applications
|July 2, 2025
Summary
This study shows that machine learning (ML) models for predicting clinical biomarkers perform better when data is stratified by sex. Analyzing data separately for males and females improves predictive accuracy in ML algorithms.
Area of Science:
- Biomedical Informatics
- Machine Learning in Healthcare
- Clinical Biomarker Discovery
Background:
- Biomarkers are essential for disease diagnosis, prognosis, and treatment selection.
- Machine learning (ML) offers powerful tools for identifying novel biomarkers and improving predictive models.
- A significant concern in ML is the potential for sex-based bias, impacting model generalizability and fairness.
Purpose of the Study:
- To develop a supervised ML model for predicting nine common clinical biomarkers.
- To investigate the impact of sex-based data stratification on ML model performance.
- To assess and mitigate sex-based bias in ML algorithms for biomarker prediction.
Main Methods:
- Development of a supervised ML model to predict triglycerides, BMI, waist circumference, systolic blood pressure, blood glucose, uric acid, urinary albumin-to-creatinine ratio, high-density lipoproteins, and albuminuria.
- Evaluation of model performance using prediction error rates (within 5-10% of actual values).
- Comparative analysis of model performance using sex-stratified data versus combined data (with and without sex as a feature).
Main Results:
- The ML model achieved prediction accuracy within 5-10% error for the selected clinical biomarkers.
- Top-performing models for predictions within 10% error included waist circumference, albuminuria, BMI, blood glucose, and systolic blood pressure.
- Sex-stratified analysis showed superior performance compared to combined datasets; models using combined data without sex as a feature performed poorest.
Conclusions:
- Stratifying data by sex significantly benefits the performance of ML-based models for clinical biomarker prediction.
- Addressing sex as a biological variable is crucial for developing accurate and unbiased ML algorithms in healthcare.
- This approach enhances the reliability of ML models for clinical decision-making and personalized medicine.

