Rapid assessment of the intrinsic hazard potential of different types of plastic additives using machine learning
Qiang Liu1, Ye Zhang1, Zhen Hua Liu1
1College of Engineering, Jilin Normal University, Siping, Jilin 136000, China.
Abstract:
As plastic pollution becomes increasingly severe, the additives released during its degradation have attracted widespread attention. How to effectively assess the potential hazards of plastic additives and identify candidates with relatively lower predicted intrinsic hazard potential for further evaluation has become an urgent issue. Experimental hazard assessment is costly and time-consuming, and the limited availability of baseline data prevents comprehensive screening of the large number of plastic additives in use. Currently, most additives lack sufficient baseline data, making it difficult for existing risk assessment methods to provide effective predictions under conditions of data scarcity. This paper proposes an intrinsic hazard potential assessment method based on the random forest-multilayer perceptron (RF-MLP) model to address this challenge. By combining the powerful nonlinear modeling capabilities of RF with the deep feature learning capabilities of Multi-Layer Perceptrons, this method can effectively capture the complex relationships between plastic additives and intrinsic hazard potential, thereby improving the accuracy of intrinsic hazard potential predictions. To further enhance the model's robustness and predictive capability, an Out-of-Fold (OOF) combination strategy was incorporated into the model. Linear regression analysis was performed on the outputs of the RF and MLP, thereby improving the accuracy of intrinsic hazard potential predictions. This model was used to conduct intrinsic hazard potential assessments on 7267 plastic additives. The RF-MLP model achieved an R2 of 0.68 on the independent test set, reflecting the degree of agreement between the predicted and observed Hazard_score_sum values. Additionally, the performance of the RF-MLP model was compared with that of commonly used machine learning models (such as RF, MLP, SVR, and XGBoost) and other hybrid models. Validation was conducted on 60 plastic additives across different predicted intrinsic hazard potential intervals, incorporating toxicity or intrinsic hazard potential results reported in the literature and data from the T.E.S.T. software developed by the U.S. EPA. The results indicate that the RF-MLP model outperforms the comparison models in predictive performance. Comparisons with literature reports and T.E.S.T. outputs provided complementary evidence for interpreting the predicted hazard patterns. The constructed RF-MLP model partially mitigates the practical limitations caused by the prediction challenges caused by insufficient baseline data and can be applied not only to the intrinsic hazard potential assessment of plastic additives but also serve as a reference for chemical intrinsic hazard potential prediction in other data-scarce scenarios.
