Related Experiment Video
Updated: Jul 16, 2026

07:18
Evaluation of Capillary and Other Vessel Contribution to Macular Perfusion Density Measured with Optical Coherence Tomography Angiography
Published on: February 18, 2022
1.9K
CoT-XNet: contextual transformer with Xception network for diabetic retinopathy grading
Shuiqing Zhao1,2, Yanan Wu1, Mengmeng Tong3
1College of Medicine and Biological Information Engineering, Northeastern University, Shenyang, People's Republic of China.
Physics in Medicine and Biology
|November 2, 2022
Summary
A new AI model, CoT-XNet, improves diabetic retinopathy (DR) grading by combining vision transformers and Xception networks. This approach enhances diagnostic accuracy for DR screening in large populations.
Area of Science:
- Ophthalmology
- Medical Imaging
- Artificial Intelligence
Background:
- Diabetic retinopathy (DR) grading relies on fundus image analysis, but subtle lesions pose challenges for current deep learning models like CNNs.
- Vision transformers offer alternative visual representations, potentially improving upon CNN performance in medical image analysis.
Purpose of the Study:
- To develop and evaluate a novel deep learning model, CoT-XNet, for enhanced accuracy in diabetic retinopathy grading.
- To leverage the complementary strengths of vision transformers and Xception networks for improved DR detection.
Main Methods:
- Proposed a two-path CoT-XNet model integrating contextual transformer (CoT) and Xception network representations.
- Implemented dedicated pre-processing, data resampling, and test-time augmentation strategies.
- Evaluated performance on over 50,000 images across three public datasets (DDR, APTOS2019, EyePACS) and compared with SOTA models.
Main Results:
- CoT-XNet outperformed existing state-of-the-art models in DR grading accuracy and Kappa scores across all tested datasets.
- Achieved accuracies of 83.10%, 84.18%, and 84.10% with Kappa values of 0.8496, 0.9000, and 0.7684, respectively.
- Class activation maps indicated that CoT and Xception networks provide different and complementary visual insights.
Conclusions:
- The CoT-XNet model demonstrates superior performance and generalizability in grading diabetic retinopathy from fundus images.
- Integrating diverse visual representations from CoT and Xception networks enhances diagnostic accuracy.
- CoT-XNet shows promise for widespread application in AI-assisted DR screening programs.

