Related Experiment Videos
Integrating Environmental Exposure Profiles with Temporal Transcriptomics: An Explainable Machine Learning Model for
Ning Sun1, Wen-Qiang Cui2, Xiang-Qing Xu2
1Acupuncture and Tuina School, Chengdu University of Traditional Chinese Medicine, Chengdu, Sichuan, 611137, China.
Background:
Cardiovascular disease (CVD) pathogenesis is strongly associated with environmental exposures and metabolic factors. This study aimed to investigate the relationship between urinary levels of volatile organic compounds (VOCs), heavy metals, and serum biomarkers with CVD, and to develop a high-accuracy, interpretable predictive model, while elucidating the potential underlying biological mechanisms from a molecular temporal-dynamic perspective.
Methods:
This cross-sectional study utilized data from the National Health and Nutrition Examination Survey (NHANES, 2011-2020 cycles), including 8,165 eligible participants. Sixty-six metabolic-related features were systematically selected from blood and urine samples. To address class imbalance, and given that the primary objective of this study was to improve the identification of cardiovascular disease (CVD) cases rather than to estimate the population prevalence distribution, the Synthetic Minority Over-sampling Technique (SMOTE) was applied exclusively to the training set for class balancing during model development. NHANES survey weights were retained for descriptive analyses and population characteristic estimation to preserve the representativeness of the survey data and minimize potential impacts on population-level inference. Least Absolute Shrinkage and Selection Operator (LASSO) regression identified the top 20 features most predictive of CVD. Seven machine learning models-including XGBoost, LightGBM, Support Vector Machine (SVM), and others-were constructed. Model performance was comprehensively evaluated using the area under the receiver operating characteristic curve (AUC-ROC), area under the precision-recall curve (AUC-PR), sensitivity, specificity, F1-score, and other metrics. The optimal model was interpreted using SHapley Additive exPlanations (SHAP) values to assess feature importance. To further validate the biological plausibility of the identified predictors, we integrated clinical transcriptomic data and conducted longitudinal analyses of differential gene expression and pathway enrichment across four critical time points (acute phase: 1 day; subacute phase: 4-6 days; recovery phase: 1 month; chronic phase: 6 months).
Results:
Among all models, XGBoost demonstrated superior predictive performance and stability, achieving an AUC of 0.845 and an accuracy of 78.24%. The 20 key features identified by LASSO regression included urinary VOC-related metabolites (e.g., ATCA and NAE), urinary inorganic arsenic species As (III), metal/element biomarkers (e.g., iodine, cadmium, and cesium), and serum biomarkers such as creatinine, glycated hemoglobin (HbA1c), and alkaline phosphatase (ALP). SHAP analysis quantified the contribution of these features to CVD risk and effectively stratified individuals into high- and low-risk groups, suggesting that the model has reasonable risk-stratification ability and interpretability. Furthermore, clinical transcriptomic time-series analysis provided molecular mechanistic support for the key predictive features: the predictive value of monocyte count was associated with significant upregulation of core monocyte/macrophage genes (e.g., CD14, S100A9) during the acute phase, while the predictive power of HbA1c corresponded with sustained dysregulation of metabolism-related genes (e.g., HBB) in the chronic phase. These findings validate the biological plausibility of the predictive model at the molecular level.
Conclusion:
Integrating urinary environmental exposures (VOCs and heavy metals) with blood biomarkers significantly enhances the accuracy of CVD risk prediction. The XGBoost-based framework, interpreted via SHAP, provides an interpretable approach for identifying key environmental and metabolic factors contributed to CVD risk prediction and exploring potential nonlinear patterns among multidimensional exposures. Multi-omics integrative analysis further elucidated, from a molecular temporal-dynamic perspective, the coherent pathophysiological mechanisms underlying these predictive features-spanning from innate immune activation, molecular patterns consistent with metabolic memory-like dysregulation, and environmental exposures-thereby providing exploratory support for the biological plausibility of the model. Future prospective cohort studies are warranted to validate the causal pathways linking these biomarkers to CVD.