Related Experiment Video
Updated: Sep 11, 2026

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
ARTNet: Adaptive channel-wise and Region-aware Transformer Network for diabetic retinopathy segmentation and
Annuj Kumar1, Munjam Shruthi1, Thiruppathy Sujeeth1
1School of Computer Science and Engineering, Vellore Institute of Technology, Chennai, India.
Abstract:
Early and accurate detection of diabetic retinopathy (DR) is essential to prevent irreversible vision loss; however, manual screening is labor-intensive and subject to inter-observer variability. To address these limitations, we propose ARTNet, an Adaptive channel-wise and Region-aware Transformer Network for automated DR classification and segmentation from retinal fundus images. ARTNet integrates three sub-network mechanisms. The Adaptive Channel-wise Feature Network (ACFNet) performs channel recalibration using dual pooling and shared multilayer perceptrons to enhance discriminative retinal representations while suppressing irrelevant responses. The Ophthalmic Region-Aware Attention Network (ORAANet) applies spatial attention to highlight clinically significant regions, including lesions and abnormal vasculature. The Retinal Patch Aggregation Encoder Network (RPAENet), built on multi-head self-attention, captures long-range dependencies and global retinal context for hierarchical feature modeling. Convolutional refinement and global average pooling enable robust five-class DR classification, while a class-balanced focal loss mitigates data imbalance and improves minority-class sensitivity. Extensive experiments on benchmark datasets demonstrate the superiority of ARTNet over intermediate and state-of-the-art models. On the Diabetic Retinopathy Detection dataset, ARTNet achieves 96.45% accuracy, 96.85% precision, 96.34% recall, and 96.60% F1-score. The model is further validated on the APTOS-2019 Blindness Detection and IDRiD datasets. Classification performance is evaluated using image-level DR grading metrics, whereas lesion segmentation performance is evaluated using the pixel-level lesion annotations available only in the IDRiD dataset. Results show that dual attention with transformer-based global reasoning improves feature representation and classification reliability. Its efficiency and stability support real-time ophthalmic screening and a clinical decision system.
