使用多类ARKA框架和基于机器学习的堆叠回归来预测有机化学品致癌性的新方法方法 (NAM)
Arkaprava Banerjee1, Kunal Roy1
1Drug Theoretics and Cheminformatics Laboratory, Department of Pharmaceutical Technology, Jadavpur University, Kolkata 700 032, India.
Journal of hazardous materials
|July 25, 2025
概括
本研究引入了新的定量结构-活性关系 (QSAR) 模型,以预测有机污染物的致癌风险. 开发的模型准确地估计了口腔斜率因子 (OSF) 和吸入斜率因子 (ISF),有助于环境风险评估.
科学领域:
- 环境化学环境化学
- 毒理学 毒理学 毒理学
- 计算化学是一种计算化学.
背景情况:
- 有机污染物带来了重大的环境和健康风险,包括致癌性.
- 准确预测致癌潜力对于风险评估和管理至关重要.
- 预测致癌性的现有方法可能是复杂和耗时的.
研究的目的:
- 开发和验证定量结构-活动关系 (QSAR) 模型,用于预测有机污染物的口腔倾斜因子 (OSF) 和吸入倾斜因子 (ISF).
- 通过结合交叉读取衍生相似度 (定量交叉读取结构-活动关系 - q-RASAR) 和特征贡献分析 (K组算术余量分析 - ARKA) 来增强模型性能.
- 根据其致癌风险识别和优先考虑有机污染物.
主要方法:
- 使用各种基于机器学习的堆叠回归器开发QSAR模型.
- 整合q-RASAR用于相似性测量和ARKA用于特征分析.
- 使用多标准决策方法选择表现最佳的模型.
- 使用外部数据集验证模型.
主要成果:
- 线性支向量回归模型在OSF预测方面取得了最佳表现 (MAE_Test=0.907).
- 斜坡回归模型在ISF预测方面表现出卓越的性能 (MAE_Test=0.827).
- 改进的建模工作流程显著提高了预测准确性.
- 对外部数据集的预测显示出与报告的致癌性状况有很好的一致性.
结论:
- 开发的QSAR模型为OSF和ISF提供了可靠的预测,有助于评估有机污染物的致癌性.
- 整合q-RASAR和ARKA框架提高了QSAR模型的预测能力.
- 这些模型可以成为优先考虑环境污染物的有价值工具,并减轻健康风险.
更多相关视频
07:35Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
7.6K
05:47In Silico Modeling Method for Computational Aquatic Toxicology of Endocrine Disruptors: A Software-Based Approach Using QSAR Toolbox
Published on: August 28, 2019
14.1K
相关概念视频
Mutagenicity and Carcinogenicity
1.4K
Mutagenicity and carcinogenicity refer to the ability of drugs to cause genetic defects and induce cancer, respectively. The International Agency for Research on Cancer (IARC) classifies agents into four groups based on their carcinogenic potential. Group 1 agents are known human carcinogens; group 2A agents are probably carcinogenic to humans; group 3 agents lack data to support their role in carcinogenesis; and group 4 includes agents for which data support that they are not likely to be...
1.4K
Mechanistic Models: Compartment Models in Individual and Population Analysis
87
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
87
