Interpretable Machine Learning with SHAP Identifies Key Biomarkers in a Multi-Factorial Spectrum of Age-Related
Daniil V Artamonov1,2, Polina I Popova3, Ekaterina A Korf4
1Group of Theoretical Chemistry, N.D. Zelinsky Institute of Organic Chemistry of Russian Academy of Sciences, Leninsky Prospect 47, Moscow 119991, Russia.
International Journal of Molecular Sciences
|February 27, 2026
Summary
This study highlights how variance-aware machine learning improves diagnosis of elderly vascular and metabolic disorders. Interpretable models identify key biomarkers like iron and glucose for better diagnostic accuracy.
Area of Science:
- Biomedical data analysis
- Geriatric medicine
- Machine learning in healthcare
Background:
- Vascular and metabolic disorders in the elderly, including stroke and diabetes, are difficult to diagnose using standard blood tests.
- Multi-parameter blood chemistry analysis is crucial but complex for these conditions.
- Heterogeneity of variance in biochemical features can distort traditional statistical analyses.
Purpose of the Study:
- To evaluate machine learning models for diagnosing elderly vascular and metabolic disorders.
- To identify key biochemical biomarkers and their interactions using interpretable AI.
- To assess the impact of variance-aware statistical methods on diagnostic accuracy.
Main Methods:
- Analyzed 49 biochemical features in 120 elderly patients with vascular/metabolic disorders.
- Applied variance-aware statistical testing to identify features with heterogeneous variance.
- Developed and compared standard ML classifiers, a gradient boosting model (max depth=3), and KMeans clustering.
- Utilized Shapley Additive Explanations (SHAP) for model interpretability.
Main Results:
- Gradient boosting model achieved high discriminative accuracy (F1-scores 0.87-0.96).
- Key biomarkers identified by SHAP include iron (Fe), transferrin, and glucose, showing synergistic interactions.
- Variance-aware testing revealed significant heterogeneity in features like Fe, Transf, RDW%, and LDL.
- Unsupervised KMeans clustering showed poor agreement with clinical labels (ARI=0.198, NMI=0.286).
Conclusions:
- Interpretable machine learning, particularly gradient boosting with restricted depth, enhances diagnostic reliability for elderly vascular and metabolic disorders.
- Variance-aware statistical approaches are crucial for handling biochemical data heterogeneity.
- Biomarker interactions, such as those between iron, transferrin, and glucose, are important for accurate diagnosis.


