Related Experiment Video
Updated: Mar 27, 2026

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
A unified VQ-VAE framework for few-shot retinal vessel segmentation and multidisease classification
Haojun Yu1, Zongcai Tan2, Huazhen Liu3
1Clinical Medicine, Chongqing Medical University-University of Leicester Joint Institute, Chongqing Medical University, Chongqing, China.
Purpose:
To address the reliance of task-specific deep learning models on large annotated datasets, this study investigates a Vector Quantized Variational Autoencoder (VQ-VAE) based few-shot learning framework for retinal vessel segmentation and disease classification.
Methods:
A compact VQ-VAE was pretrained on unlabeled fundus photographs to learn transferable discrete representations. The pretrained encoder was used to initialize multiple downstream models, including segmentation networks (U-Net, SegNet, ERFNet) and three additional architectures (FR-UNet, Swin-Res-Net, RV-GAN), as well as classification networks (VGG-16, ResNet-50, EfficientNet-B0). Retinal vessel segmentation was evaluated on three public datasets (DRIVE, STructured Analysis of the Retina [STARE], and CHASE), while disease classification was assessed on the Retina and Ocular Disease Intelligent Recognition (ODIR) datasets. Segmentation performance was evaluated using Dice coefficient, Recall, Accuracy, Intersection over Union (IoU), mean IoU, and Average Offset Distance. Classification performance was assessed using accuracy, precision, recall, F1-score, and area under the receiver operating characteristic curve (AUC), with additional validation on the heterogeneous ODIR dataset.
Results:
The VQ-VAE pretraining consistently improved performance under small-sample conditions. On the DRIVE dataset, Dice scores increased by approximately 2 percentage points across all segmentation backbones, with U-Net improving from 0.780 to 0.796 and SegNet from 0.670 to 0.692. Consistent performance improvements were also observed on the STARE and CHASE datasets. For disease classification on the Retina dataset, accuracy increased from 20% to 60% using only 70 labeled images. On the ODIR dataset, mean AUC improved across architectures, from 0.677 to 0.724 for VGG-16, 0.684 to 0.745 for ResNet-50, and 0.664 to 0.726 for EfficientNet-B0, indicating enhanced robustness across diverse disease categories and imaging conditions.
Conclusion:
The proposed pretraining framework effectively reduces labeled data requirements while improving performance across multiple ophthalmic tasks, offering a scalable and resource-efficient solution for real-world clinical applications.

