YOLOv5 Attention Analysis for Anterior Eye Disease Classification: Grad-CAM++ Feature Importance and Cut-and-Paste
Yoshiyuki Kitaguchi1,2, Yuta Ueno3,4,5, Takefumi Yamaguchi6,4
1Department of Ophthalmology, Osaka University Graduate School of Medicine, 2-2 Yamadaoka, Suita, Osaka, 565-0871, Japan. kitaguchi@ophthal.med.osaka-u.ac.jp.
Journal of Imaging Informatics in Medicine
|October 14, 2025
Summary
This study enhances AI diagnostics for eye diseases by combining Grad-CAM++ and cut-and-paste validation. It reveals how AI models use background context, improving diagnostic transparency for anterior segment diseases.
Area of Science:
- Ophthalmology
- Artificial Intelligence
- Medical Imaging
Background:
- Deep learning models like YOLOv5 are increasingly used for diagnosing anterior segment diseases from slit-lamp images.
- Improving the explainability of these AI models is crucial for clinical trust and adoption.
- Understanding the influence of contextual information on diagnostic accuracy is essential for robust AI systems.
Purpose of the Study:
- To enhance the explainability of a YOLOv5 model for anterior segment disease diagnosis.
- To combine Gradient-weighted Class Activation Mapping++ (Grad-CAM++) with cut-and-paste validation for model interpretability.
- To evaluate the impact of extracorneal (background) information on the diagnostic accuracy of the AI model.
Main Methods:
- A YOLOv5 model was trained and evaluated on 1039 slit-lamp photographs across nine diagnostic categories.
- Grad-CAM++ was used to visualize model attention patterns across network layers.
- Cut-and-paste validation was performed by altering image backgrounds to assess reliance on contextual information.
- Intersection over Union (IoU) was used to compare attention maps with expert annotations.
Main Results:
- Grad-CAM++ showed hierarchical attention, with layer 23 exhibiting disease-specific patterns.
- Model accuracy significantly decreased for certain diseases (e.g., infectious keratitis, acute primary angle closure) when backgrounds were altered.
- Attention maps had higher clinical relevance (IoU) for correct predictions, linking visual explanations to diagnostic reliability.
Conclusions:
- The combination of Grad-CAM++ and cut-and-paste validation offers a robust method for assessing AI model explainability in ophthalmology.
- This approach elucidates layer-specific attention mechanisms and quantifies the model's dependence on extracorneal features.
- The findings contribute to greater transparency and trustworthiness in AI-driven diagnostic systems for eye care.


