Related Experiment Video
Updated: Jun 29, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Using test-time augmentation to investigate explainable AI: inconsistencies between method, model and human
Peter B R Hartog1,2, Fabian Krüger3, Samuel Genheden4
1Molecular AI, Discovery Sciences, R &D, AstraZeneca, 431 83, Mölndal, Sweden. peter.hartog@astrazeneca.com.
Explainable artificial intelligence (XAI) methods show inconsistencies for molecular representations in computational toxicity. Test-time augmentation reveals that explanations may reflect tokenization rather than learned parameters, urging caution in model validation.
Area of Science:
- Computational toxicology
- Cheminformatics
- Artificial intelligence
Background:
- Machine learning models require explainable artificial intelligence (XAI) for human-understandable interpretations.
- Text-based molecular representations are crucial for transfer learning in computational toxicity.
- Augmenting molecular representations aids in comparing model outputs for the same data.
Purpose of the Study:
- To investigate the robustness of eight XAI methods using test-time augmentation for molecular representation models in computational toxicity prediction.
- To assess the consistency and reliability of XAI explanations for identical molecular structures under different representations.
Main Methods:
- Utilized test-time augmentation on text-based molecular representations.
- Evaluated eight different explainable artificial intelligence (XAI) methods.
- Compared explanations generated for multiple representations of the same molecular structure.
- Analyzed the variance between in-domain and out-of-domain predictions.
- Assessed the importance assigned to expert-derived structural alerts.
Main Results:
- Significant differences were observed in explanations generated for different representations of the same molecule.
- Randomized models exhibited similar variance in explanations compared to standard models.
- XAI measures showed greater variance for in-domain predictions than out-of-domain predictions.
- Expert-derived structural alerts received similar importance across various conditions, irrespective of applicability domain or model randomization.
- Inconsistencies were found within models for identical molecular representations.
Conclusions:
- Current XAI methods may not reliably reflect learned parameters in text-based molecular representations, potentially indicating a reliance on tokenization.
- Test-time augmentation is a valuable tool for assessing the consistency of XAI in computational toxicology.
- Researchers should validate XAI methods by comparing them to human intuition and expert knowledge, especially when using text-based molecular representations.
More Related Videos
14:14The Innovation Arena: A Method for Comparing Innovative Problem-Solving Across Groups
Published on: May 13, 2022
05:22Dissociation of the Confounding Influences of Expectancy and Integrative Difficulty Residing in Anomalous Sentences in Event-related Potential Studies
Published on: May 9, 2019
Related Concept Videos
Randomized Experiments
Simple randomization
Simple...
Improving Translational Accuracy
Generalization, Discrimination, and Extinction
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
Fundamental Attribution Error
Introduction to Test of Independence
The test statistic for a test of independence is similar to that of a goodness-of-fit test: