使用测试时间增强来调查可解释的AI:方法,模型和人类直觉之间的不一致性
Peter B R Hartog1,2, Fabian Krüger3, Samuel Genheden4
1Molecular AI, Discovery Sciences, R &D, AstraZeneca, 431 83, Mölndal, Sweden. peter.hartog@astrazeneca.com.
可解释的人工智能 (XAI) 方法在计算毒性中显示分子表示的不一致性. 测试时间增长显示,解释可能反映了代币化,而不是学习的参数,敦促在模型验证中谨慎行事.
科学领域:
- 计算毒理学计算毒理学
- 化学信息学 化学信息学
- 人工智能的人工智能是人工智能.
背景情况:
- 机器学习模型需要可解释的人工智能 (XAI) 来实现人类可理解的解释.
- 基于文本的分子表示对于计算毒性转移学习至关重要.
- 增强分子表示有助于对同一数据的模型输出进行比较.
研究的目的:
- 在计算毒性预测中研究八种XAI方法的稳定性,使用测试时间增强用于分子表示模型.
- 评估XAI解释在不同的表示下对相同分子结构的一致性和可靠性.
主要方法:
- 在基于文本的分子表示上利用了测试时间的增强.
- 评估了八种不同的可解释的人工智能 (XAI) 方法.
- 对同一个分子结构的多个表示生成的比较解释.
- 分析了域内和域外预测之间的差异.
- 评估了专家衍生结构性警报所赋予的重要性.
主要成果:
- 对于相同分子的不同表示生成的解释中观察到显著差异.
- 与标准模型相比,随机模型在解释中表现出类似的差异.
- XAI测量显示,域内预测的变异性比域外预测的变异性更大.
- 来自专家的结构性警报在各种条件中都具有相似的重要性,无论适用领域或模型随机化.
- 在同样的分子表示模型中发现了不一致性.
结论:
- 当前的XAI方法可能无法在基于文本的分子表示中可靠地反映学习的参数,这可能表明它依赖于令牌化.
- 测试时间增长是评估XAI在计算毒理学中的一致性的一个有价值的工具.
- 研究人员应该通过将它们与人类直觉和专家知识进行比较来验证XAI方法,特别是当使用基于文本的分子表示时.
更多相关视频
14:14The Innovation Arena: A Method for Comparing Innovative Problem-Solving Across Groups
Published on: May 13, 2022
05:22Dissociation of the Confounding Influences of Expectancy and Integrative Difficulty Residing in Anomalous Sentences in Event-related Potential Studies
Published on: May 9, 2019
相关概念视频
Randomized Experiments
Simple randomization
Simple...
Improving Translational Accuracy
Generalization, Discrimination, and Extinction
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
Fundamental Attribution Error
Introduction to Test of Independence
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
