改善液体染色学质谱学化合物适应性的预测,以增强非向分析
Nathaniel Charest1, Charles N Lowe2, Christian Ramsland3
1Center for Computational Toxicology and Exposure, Office of Research and Development, U.S. Environmental Protection Agency, Research Triangle Park, North Carolina, 27711, USA. charest.nathaniel@epa.gov.
这项研究增强了在毒理风险评估中用于非定向分析 (NTA) 的机器学习模型. 改进的模型更好地预测化学物质对质谱学的适应性,有助于污染物识别和风险表征.
科学领域:
- 环境化学环境化学
- 分析化学 分析化学
- 毒理学 毒理学 毒理学
背景情况:
- 使用质谱的非向分析 (NTA) 对于识别环境污染物至关重要.
- 在NTA中准确的化学识别依赖于预测化合物对质谱法方法的适应性.
- 以前的模型预测了化学易受性,但解释性和数据范围需要改进.
研究的目的:
- 提高机器学习模型的可解释性和性能,用于预测液态染色体质谱学 (LC-MS) 的化学易受性.
- 将数据集扩展为正负电离模式的新型策划数据.
- 通过可解释的机器学习增强对NTA识别的信心.
主要方法:
- 使用特征工程开发和完善随机森林模型,以提高可解释性.
- 整合了来自专家策划的化学化合物标签的1348个额外数据点.
- 在内部和外部数据集上使用平衡精度 (BA) 和马修斯相关系数 (MCC) 评估模型性能.
主要成果:
- 新型模型显示了与先前工作相比的平衡精度 (平均CV BA 0.84 / 0.85与0.82).
- 与之前的模型 (0.55/0.54) 相比,在外部数据集上的马修斯相关系数 (MCC) 显著改善 (0.66/0.68).
- 增强的数据集和特征工程使得可以更好地预测整个电离模式的化学易受性.
结论:
- 增强的,可解释的机器学习模型为NTA在毒理风险评估中提供了更强大的工具.
- 改进的可接受性预测扩大了NTA在环境污染物识别中的适用性领域和可靠性.
- 这项工作支持正在进行的努力,以开发高性能,可解释的模型来推进NTA能力.
更多相关视频
07:34Large Scale Non-targeted Metabolomic Profiling of Serum by Ultra Performance Liquid Chromatography-Mass Spectrometry UPLC-MS
Published on: March 14, 2013
09:04Identifying Per- and Polyfluorinated Chemical Species with a Combined Targeted and Non-Targeted-Screening High-Resolution Mass Spectrometry Workflow
Published on: April 18, 2019
相关概念视频
Mass Spectrometry: Complex Analysis
GC–MS is a powerful hyphenated method commonly used in forensics and environmental...
High-Performance Liquid Chromatography: Types of Detectors
Peptide Identification Using Tandem Mass Spectrometry
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
