Related Experiment Video
Updated: May 14, 2026

Evaluation of Capillary and Other Vessel Contribution to Macular Perfusion Density Measured with Optical Coherence Tomography Angiography
Published on: February 18, 2022
Performance of Artificial Intelligence Systems for Automated Segmentation and Quantification of Retinal Fluid and
Bakhtawar Awan1, Mohamed Elsaigh2, Mohamed Hesham Gamal3,4
1General Surgery, Northwick Park Hospital, London, GBR.
Abstract:
Optical coherence tomography (OCT) plays a crucial role in diagnosing retinal diseases, such as diabetic retinopathy (DR) and age-related macular degeneration (AMD), as well as in identifying neurodegenerative biomarkers. Despite advancements in U-Net-based convolutional networks for OCT image segmentation, there is a lack of systematic reviews comparing their performance with expert manual segmentations. This review aims to assess the efficacy of these automated networks in segmenting retinal fluid and pathology in OCT images. By searching three different databases, PubMed, Web of Science, and Scopus, over the past five years, we conducted this systematic review and meta-analysis using data from 16 diagnostic-accuracy studies. Study quality was assessed using the Quality Assessment of Diagnostic Accuracy Studies-2 (QUADAS-2) tool. The analysis used mean and standard deviation for the continuous outcomes and employed a random-effects model. Analyses were performed using Review Manager software version 5.4 (The Cochrane Collaboration, London, UK, 2020). Artificial intelligence (AI) and human Dice scores did not differ significantly (standardized mean difference (SMD) = -0.08; 95% CI: -1.16 to 0.99; p = 0.88), nor did intraclass correlation coefficient (ICC) values (SMD = -0.13; 95% CI: -5.70 to 5.45; p = 0.96). However, very high heterogeneity (I² > 90%) limits the reliability of these pooled estimates. AI achieved expert-level Dice scores for subretinal fluid (0.88-0.96) and geographic atrophy (0.94). Intraretinal fluid was more challenging (Dice 0.79-0.89). Volumetric reliability was strong (ICCs > 0.94). Device-dependent variability was substantial; kappa was 0.37 for ZEISS versus 0.73 for Spectralis, indicating a need for device-specific optimization. Volumetric analyses revealed minor systematic overestimation (mean difference: -0.05 mm²). Processing times ranged from 100 milliseconds per B-scan to several seconds per volume, representing substantial time savings versus manual segmentation. Fully automated U-Net pipelines reach expert-level accuracy for subretinal fluid and geographic atrophy but remain limited for intraretinal fluid and show marked device-dependent variability. Clinical translation requires four priorities: standardized multi-device benchmarks, domain adaptation for cross-platform robustness, hybrid AI-human workflows pairing automated pre-segmentation with expert oversight, and prospective clinical trials. These steps are needed to move AI segmentation from a research tool to a clinical decision-support system.
