Related Experiment Videos
Code-Free Automated Machine Learning for Referable Glaucoma Classification From Retinal Fundus Photographs: A
Mark A Bachir1, Alexander Bachir2, Neel Nawathey3
1Internal Medicine, California Northstate University College of Medicine, Elk Grove, USA.
Abstract:
Artificial intelligence has demonstrated substantial potential for automated glaucoma screening from retinal fundus photographs; however, conventional deep-learning development often requires programming expertise, specialized knowledge of neural-network design, and substantial computational resources. This study evaluated whether a commercially available, code-free automated machine-learning (AutoML) platform could generate a high-performing classifier for distinguishing referable glaucoma (RG) from non-referable glaucoma (NRG). A publicly available balanced retinal fundus image dataset derived from the EyePACS Artificial Intelligence for Robust Glaucoma Screening dataset was used. The study included 9,540 color fundus photographs, with 4,770 images in each class. Images were assigned to training (n=7,632; 80%), validation (n=954; 10%), and held-out internal testing (n=954; 10%) partitions. A single-label image-classification model was developed using Google Cloud Vertex AI AutoML (Google LLC, California, US) through its graphical interface without investigator-written machine-learning code or manual neural-network architecture design. Model training used a maximum computational budget of eight node-hours. Performance was evaluated using average precision, precision, recall, precision-recall analysis, confidence-threshold analysis, and a row-normalized confusion matrix. The AutoML model achieved an overall average precision of 0.974. At the default confidence threshold of 0.50, overall precision and recall were both 92.5%. Class-specific average precision was 0.977 for both RG and NRG. The model correctly classified 94% of NRG images and 91% of RG images. Deployment to a cloud-based prediction endpoint successfully generated class predictions and confidence scores for representative retinal fundus photographs. A fully graphical AutoML workflow therefore produced strong internal classification performance for distinguishing RG from NRG without investigator-written machine-learning code or manual neural-network development. These findings support the feasibility of commercial AutoML platforms as accessible tools for ophthalmic image-classification research. Because the model was evaluated using a single derived dataset without external testing, further validation across independent populations, institutions, and imaging systems is necessary before clinical implementation.