Related Experiment Video
Updated: Oct 10, 2026

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
SEFCA-Net: a semantic-edge-frequency cross-attention network for diabetic retinopathy lesion segmentation
Jin Xie1, Jingtao Wang2, Duo Xu3
1Nanning Normal University, Guangxi Key Lab of Human-Machine Interaction and Intelligent Decision, Nanning, China.
Significance:
Diabetic retinopathy (DR) lesion segmentation provides pixel-level lesion localization in fundus images. However, such lesions are often tiny, sparsely distributed, low-contrast, morphologically diverse, and visually similar to vessels or surrounding retinal structures, making accurate localization difficult. Many existing methods emphasize multi-scale spatial representation learning and handle boundary ambiguity and subtle texture differences mainly through encoder-decoder fusion. We therefore examine boundary and frequency representations as explicit, complementary inputs to semantic decoding.
Aim:
We aim to develop SEFCA-Net, a semantic-edge-frequency cross-attention network for diabetic retinopathy lesion segmentation. The main objective is to examine whether semantic, edge, and frequency-domain representations can be fused under semantic guidance to improve pixel-level lesion localization without compromising the semantic decoding stream. The evaluation focuses on four lesion categories, including hard exudates, hemorrhages, soft exudates, and microaneurysms, using public benchmarks and a DDR-to-IDRiD cross-dataset setting.
Approach:
SEFCA-Net uses HRNet-W48 as the semantic backbone to extract multi-scale semantic features and explicitly constructs edge and frequency-domain representations from input fundus images. The learnable multi-scale edge branch extracts boundary-sensitive features, while the all-channel multi-scale discrete cosine transform (DCT) frequency-domain branch generates texture-sensitive responses to low-contrast lesions. The semantic-edge-frequency cross-attention decoder uses the semantic representation as the primary stream and injects edge and frequency-domain features through semantic-guided window-based cross-attention. Performance is evaluated using the area under the precision-recall curve (AUPR), Dice, and Intersection-over-Union (IoU).
Results:
On DDR, SEFCA-Net achieves 47.85% mean AUPR (mAUPR), 46.76% mean Dice, and 31.22% mean IoU. On IDRiD, it reaches 69.74% mAUPR; when trained on DDR and tested on IDRiD, it obtains 59.83% mAUPR. Ablation studies show that either auxiliary branch improves the semantic-only baseline and that the complete configuration gives the highest mAUPR. Multi-scale Sobel modeling and DCT outperform the tested single-scale and wavelet alternatives, respectively.
Conclusions:
SEFCA-Net shows that semantic features can remain the main decoding stream while boundary-sensitive features and texture-sensitive responses are integrated in a controlled manner. Under the evaluated protocols, the model shows competitive in-domain performance and records the highest DDR-to-IDRiD mAUPR among the reproduced methods.