Related Experiment Video
Updated: Oct 5, 2026

Using Retinal Imaging to Study Dementia
Published on: November 6, 2017
External Validation and Comparison of Automated Retinal Image Quality Grading Methods in the Canadian Longitudinal
Hon Yiu So1, Tom Guo2, Sangeetha Srinivasan3
1Department of Mathematics and Statistics, Oakland University, Oakland.
Purpose:
To externally validate and compare automated retinal image quality grading using a multiple color-space fusion network (MCF-Net) and a vessel-based fractal dimension (FD) score in the Canadian Longitudinal Study on Aging (CLSA).
Design:
Cross-sectional external validation study.
Participants:
There are 7263 baseline color fundus photographs from CLSA participants.
Methods:
Images were resized and normalized and then converted to red, green, and blue/hue, saturation, and value/Lab for MCF-Net; FD was computed from the segmented vasculature. Two of 3 expert graders assigned 3-tier reference labels (good, fair, poor/ungradable). To enable fair external comparison between probabilistic (MCF-Net) and continuous (FD) outputs, predictions were mapped onto a common ordinal scale aligned with expert grading, and operating cut points were selected within the CLSA data set to address class imbalance by maximizing macro-averaged F1. Performance was also evaluated in binary tasks (good vs. non-good and poor/ungradable vs. non-poor/ungradable).
Main Outcome Measures:
Macro-averaged F1 for 3-class grading; quadratic-weighted kappa; accuracy; and F1 for poor/ungradable detection.
Results:
Multiple color-space fusion network outperformed FD on 3-class grading (accuracy, 85.0% vs. 83.9%; macro-F1, 0.739 vs. 0.673; quadratic-weighted κ, 0.843 vs. 0.792). Gains were largest for poor/ungradable detection (F1, 0.808 vs. 0.625). Agreement with expert ranking was higher for MCF-Net. Both methods performed best for good and poor/ungradable images; fair remained the most challenging category.
Conclusions:
In this population-based external validation study, MCF-Net showed modest overall gains compared with FD, but its principal advantage was improved detection of poor/ungradable images. This distinction is clinically relevant because missed poor/ungradable images may create false reassurance in screening or epidemiologic workflows. By introducing a unified evaluation framework that enables fair comparison between probabilistic and continuous outputs, this study addresses a practical barrier to validating automated image-quality tools in clinical screening and population-based cohort studies. The observed error profile of MCF-Net-marked by improved detection of poor/ungradable images and conservative reclassification of borderline cases-aligns with the priorities of screening and epidemiologic workflows, where false reassurance poses greater risk than additional image review or retakes. These findings support further development and prospective evaluation of automated retinal image quality assessment before potential integration into large-scale workflows.
Financial Disclosures:
Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.
