Related Experiment Video
Updated: Jun 7, 2025

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
Published on: April 21, 2023
GradToken: Decoupling tokens with class-aware gradient for visual explanation of Transformer network
Lin Cheng1, Yanjie Liang2, Yang Lu1
1Fujian Key Laboratory of Sensing and Computing for Smart City, School of Informatics, Xiamen University, Xiamen 361005, China.
Researchers developed GradToken, a new visual explanation method for Transformer networks. This approach effectively addresses semantic coupling in attention weights, improving the interpretability and trustworthiness of Transformer models in various applications.
Area of Science:
- Artificial Intelligence
- Computer Vision
- Machine Learning
Background:
- Transformer networks are increasingly utilized across diverse AI fields, including computer vision and natural language processing.
- Explaining Transformer decision-making is crucial for enhancing trust and understanding in real-world applications.
- Existing explanation methods for convolutional neural networks are not directly applicable to Transformer architectures due to structural differences.
Purpose of the Study:
- To address the challenge of designing effective explanation methods specific to Transformer networks.
- To overcome the semantic coupling problem in Transformer attention weight matrices that hinders category-specific explanations.
- To propose a novel method for generating visual explanations of Transformer predictions.
Main Methods:
- Analyzed the semantic coupling issue within Transformer attention weight matrices.
- Proposed GradToken, a gradient-decoupling-based token relevance method.
- Utilized class-aware gradients to decouple semantics and leveraged token relations to generate relevance maps.
Main Results:
- GradToken effectively decouples tangled semantics in class tokens, enabling category-specific explanations.
- Generated relevance maps accurately highlight target regions, improving visual explanation quality.
- Extensive experiments validated the method's effectiveness and reliability.
Conclusions:
- GradToken provides a specific and effective solution for explaining Transformer predictions visually.
- The method enhances the interpretability of Transformers by generating focused and reliable relevance maps.
- This work contributes to building greater trust and understanding in the application of Transformer models.
Related Concept Videos
Types Of Transformers
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
Energy Losses in Transformers
There are four main reasons for energy losses in transformers.
The first cause can be the high resistance of the...
Transformers in Distribution System
Distribution substation transformers come in various ratings and typically use mineral oil for insulation and cooling. To prevent moisture and air from entering the oil, some transformers use an inert gas like nitrogen to fill the...
Equivalent Circuits for Practical Transformers
In a practical transformer, each winding exhibits resistance and leakage reactance. The...
The Ideal Transformer
Ampere's Law forms the basis of understanding the magnetic field within the transformer. It states that the integral of the magnetic field intensity's...
Transformers with Off-Nominal Turns Ratios

