开发一种可解释的机器学习模型,用于预测环境相关化合物的神经毒性
Yuxing Hao1,2,3, Zhihui Duan1,4, Lizheng Liu1,3
1State Key Laboratory of Environmental Chemistry and Ecotoxicology, Research Center for Eco-Environmental Sciences, Chinese Academy of Sciences, Beijing 100085, P. R. China.
Environmental science & technology
|April 30, 2025
概括
环境污染物导致神经系统疾病. 一个新的机器学习模型准确地预测神经毒性,从人类血液样本中识别出821种潜在有害化合物.
科学领域:
- 环境健康 环境健康
- 毒理学 毒理学 毒理学
- 计算化学计算化学
- 神经科学是一个神经科学.
背景情况:
- 与环境污染物相关的神经系统疾病的患病率上升.
- 由于众多化合物和低效的传统方法,神经毒性数据存在重大差距.
- 需要对环境神经毒性进行强大和可解释的预测模型.
研究的目的:
- 开发一个高质量的,可解释的神经毒性预测模型.
- 评估各种分子表示和机器学习/深度学习算法.
- 选环境化合物对潜在的神经毒性影响.
主要方法:
- 评估了三个分子表示:指纹,描述符和图表.
- 对比了六种传统的机器学习 (ML) 算法和两种深度学习 (DL) 方法.
- 使用极端梯度提升 (XGBoost) 具有分子指纹和描述符,以获得最佳模型.
主要成果:
- 最佳的XGBoost模型实现了高训练精度 (0.93) 和AUC (0.99).
- 成功选了1170个人类血液化合物,预测了1145年的神经毒性.
- 确定了821种潜在的神经毒性化合物,其中36种具有高检测度.
结论:
- 开发了一种有效和可解释的模型来预测环境神经毒性.
- 该模型有助于识别高风险化合物和管理环境健康.
- 有一个在线平台可用于提高研究人员和公共卫生官员的可访问性.
更多相关视频
09:01A High-throughput Assay for the Prediction of Chemical Toxicity by Automated Phenotypic Profiling of Caenorhabditis elegans
Published on: March 14, 2019
7.2K
07:41A Neurite Outgrowth Assay and Neurotoxicity Assessment with Human Neural Progenitor Cell-Derived Neurons
Published on: August 6, 2020
7.5K
相关概念视频
Types of Toxins
1.7K
Humans continually engage with an environment rich in potentially harmful chemicals. These are introduced to our bodies through inhalation, ingestion, or skin contact. These chemicals exist in various forms, such as air and environmental pollutants, agricultural chemicals, organic solvents, and heavy metals.
Air pollutants, primarily gases, pose significant threats to respiratory health, leading to conditions like hypoxia, lung cancer, and in extreme cases, death.
Environmental pollutants like...
Air pollutants, primarily gases, pose significant threats to respiratory health, leading to conditions like hypoxia, lung cancer, and in extreme cases, death.
Environmental pollutants like...
1.7K
Mechanistic Models: Compartment Models in Individual and Population Analysis
33
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
33
