Predicting the toxicity of multicomponent complex mixtures using a multi-feature fusion-based machine learning model
Yuanfan Zhao1, Renyong Jia2, Jing Zhang1
1Key Laboratory of Water Pollution Control and Wastewater Resource of Anhui Province, College of Environment and Energy Engineering, Anhui Jianzhu University, Hefei, China; Anhui Gaodi Technology Co., LTD, Luan, China.
Abstract:
Complex chemical mixtures are ubiquitous in aquatic environments posing substantial risks to aquatic organisms and human health, and the prediction of their combined toxicity is the key measure to control their ecological risks. This, however, still remains a major challenge for the two widely used classic standard additive models, concentration addition (CA) and independent action (IA), due to complex interactions, antagonism or synergism, within mixtures. Therefore, to conquer the dilemma, a novel interpretable machine learning model was constructed by integrating the predicted values by CA and IA with molecular descriptors (CIMM) against a freshwater organism Chlorella pyrenoidosa to predict the combined toxicity of complex mixture pollutants. The CIMM model stability was assessed by repeated modeling with ten random seeds, and its interpretability was examined by using mutual information, partial dependence plots and SHAP values. A dual-metric applicability domain was established based on kernel density estimation and k-nearest neighbor distance, and generalization performance was evaluated on a completely independent external validation set. The results showed that the CIMM framework achieved a mean test-set R2 of 0.9080 ± 0.0127 across ten random-seed runs, which was implemented using the seed-specific optimal Random Forest or XGBoost algorithm. Compared to the results predicted by CA alone, CIMM increased R2 by 32.2% and the RMSE decreased by 46.4%, while to those predicted by model using only molecular descriptors, the R2 improved by approximately 8.87% and the RMSE decreased by 25.38%. CA and IA predictions were the key features of CIMM that drove the model output, and they exhibited stable monotonic positive relationships with mixture toxicity. The applicability domain assessment showed that in-domain samples achieved an external validation R2 of 0.768 which was markedly higher than 0.352 for out-of-domain samples, with lower prediction uncertainty for in-domain data. This study provides a high-accuracy, interpretable and applicability-domain-constrained framework for predicting the combined toxicity of complex mixture pollutants in aquatic environment, regardless of whether toxicity interactions occur within the mixtures.
