Related Experiment Video
Updated: Jan 27, 2026

Analyzing Mitochondrial Morphology Through Simulation Supervised Learning
Published on: March 3, 2023
Enhancing accuracy and explainability in colorectal lesion classification with attention-supervised Vision
Luca Carlini1, Luca Di Stefano1, Chiara Lena1
1Dipartimento di Elettronica, Informazione, Bioingegneria (DEIB), Politecnico di Milano, Piazza Leonardo da Vinci, 32, Milano, 20133, Italy.
Supervising Vision Transformer (ViT) attention maps with expert annotations improves colorectal lesion classification accuracy and interpretability. This method enhances Paris classification performance and provides clearer, lesion-focused attention patterns for better clinical decision-making.
Area of Science:
- Artificial Intelligence in Medicine
- Computer Vision for Endoscopy
- Medical Image Analysis
Background:
- Accurate colorectal lesion assessment is crucial for treatment and cancer risk stratification.
- The Paris classification aids this assessment but faces inter-observer variability.
- Vision Transformers (ViTs) offer potential but can exhibit diffuse attention, hindering interpretability.
Purpose of the Study:
- To investigate if supervising ViT attention maps with expert annotations improves Paris classification accuracy.
- To enhance the interpretability of ViT models by focusing attention on relevant lesion regions.
- To develop a method that concurrently boosts classification performance and model explainability.
Main Methods:
- Proposed a Lesion-Focused Attention Loss (LLFA) pretraining objective.
- Used expert polyp bounding boxes to guide ViT attention to annotated lesion areas.
- Applied LLFA to six ViT architectures, followed by cross-entropy fine-tuning on the SUN dataset for Paris classification.
Main Results:
- Attention-supervised pretraining consistently improved accuracy and lesion-focused attention across ViT models.
- LLFA enhanced three-class Paris classification accuracy by up to 7 percentage points.
- LLFA outperformed a Grad-CAM consistency baseline by 5-13 percentage points, with significant association between focused attention and correct predictions.
Conclusions:
- Directly supervising ViT attention with LLFA effectively leverages expert knowledge.
- This approach jointly improves Paris classification accuracy and spatial interpretability.
- LLFA demonstrates superior performance compared to Grad-CAM-based explanation regularization.
Related Concept Videos
Vision
Improving Translational Accuracy
Improving Translational Accuracy
Uncertainty in Measurement: Accuracy and Precision
Color Vision
Bacterial Transformation
Griffith made an unexpected discovery when he killed the pathogenic strain and mixed its remains with the live, non-pathogenic strain. Not only did the mixture kill host mice, but it also contained living pathogenic bacteria that...

