Related Experiment Video
Updated: Aug 27, 2025

04:23
A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
Published on: April 21, 2023
1.9K
Focused Attention in Transformers for interpretable classification of retinal images
Clément Playout1, Renaud Duval2, Marie Carole Boucher2
1LIV4D, Polytechnique Montréal, 2500 Ch. de Polytechnique, Montréal, QC, H3T 1J4, Canada.
Medical Image Analysis
|September 23, 2022
Summary
Vision Transformers demonstrate high performance in retinal disease classification. A novel Focused Attention method improves interpretability and lesion detection compared to Convolutional Neural Networks.
Area of Science:
- Ophthalmology
- Computer Vision
- Artificial Intelligence
Background:
- Vision Transformers (ViTs) are increasingly popular for image classification due to high performance and interpretability.
- These characteristics require thorough evaluation in the context of retinal imaging.
- Current attribution methods for ViTs produce low-resolution heatmaps, limiting their clinical utility.
Purpose of the Study:
- To compare the performance of various Vision Transformers against traditional Convolutional Neural Networks (CNNs) for retinal disease classification.
- To evaluate the models' ability to handle multi-modality retinal imaging (fundus and OCT) and generalize to external datasets.
- To introduce and validate a novel mechanism, Focused Attention, for generating high-resolution, interpretable attribution maps.
Main Methods:
- Performance evaluation of multiple Vision Transformer architectures versus CNNs on retinal disease classification tasks.
- Assessment of multi-modal imaging (fundus, OCT) and external data generalization.
- Development of Focused Attention, a novel attribution method using iterative conditional patch resampling.
- Validation of interpretability and lesion detection capabilities through a survey with four retinal specialists.
Main Results:
- Vision Transformers show competitive performance in retinal disease classification.
- Focused Attention generates higher-resolution attribution maps compared to existing Transformer methods.
- Retinal specialists found Vision Transformer attribution maps more interpretable than those from CNNs.
- Focused Attention was validated as a relevant tool for lesion detection in retinal images.
Conclusions:
- Vision Transformers offer a promising alternative to CNNs for retinal disease classification, with enhanced interpretability.
- The proposed Focused Attention mechanism significantly improves the resolution and clinical relevance of attribution maps.
- This work highlights the potential of Vision Transformers and advanced interpretability techniques in ophthalmology.
Related Concept Videos
Vision
55.1K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
55.1K
Visual System
657
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
Once through the pupil, the light passes through the lens, a...
657
Anatomy of the Eyeball
7.4K
The eye is a spherical, hollow structure composed of three tissue layers. The outer layer — the fibrous tunic, comprises the sclera — a white structure — and the cornea, which is transparent. The sclera encompasses some of the ocular surface, most of which is not visible. However, the 'white of the eye' is distinctively visible in humans compared to other species. The cornea, a clear covering at the front of the eye, enables light penetration. The eye's middle...
7.4K
The Retina
69.6K
The retina is a layer of nervous tissue at the back of the eye that transduces light into neural signals. This process, called phototransduction, is carried out by rod and cone photoreceptor cells in the back of the retina.
69.6K
Association Areas of the Cortex
6.0K
Association areas are regions of the cerebral cortex that do not have a specific sensory or motor function. Instead, they integrate and interpret information from various sources to enable higher cognitive processes such as memory, learning, and decision-making. Some key association areas include the following:
Prefrontal Association Area: This area is located in the frontal lobe and is involved in planning, decision-making, and moderating social behavior. It connects with primary motor areas,...
Prefrontal Association Area: This area is located in the frontal lobe and is involved in planning, decision-making, and moderating social behavior. It connects with primary motor areas,...
6.0K

