Related Experiment Videos
Leakage-Aware Visit-Level Benchmarking Reveals Representational Overlap in Deep Learning for Bacterial Versus Fungal
1Crimson Global Academy, Level 3, Parnell, Auckland, New Zealand.
Purpose:
The purpose of this study was to benchmark deep learning for bacterial versus fungal keratitis classification from slit-lamp photographs under a leakage-aware, visit-level framework, and to characterize embedding-space geometry as a potential performance ceiling.
Methods:
This retrospective single-center study included white-light slit-lamp photographs from 101 patients with culture-confirmed bacterial or fungal keratitis (258 visits; 658 images), using strict patient-disjoint five-fold cross-validation with the clinical visit as the primary prediction unit. Image-level classifiers, multiple instance learning (MIL), and DINOv2-based retrieval were compared under three preprocessing strategies: full-frame, corneal region of interest (ROI), and lesion-centered MedSAM crops. Embedding geometry was analyzed using Uniform Manifold Approximation and Projection (UMAP). Pairwise visit-level model comparisons were performed using patient-cluster paired bootstrap on aligned out-of-fold (OOF) visit-level predictions, and probability calibration was assessed using expected calibration error (ECE) and Brier score.
Results:
At the image level, the highest area under the receiver operating characteristic curve (AUROC) was 0.691. At the visit level, convolutional neural network (CNN)-based MIL models achieved the highest AUROC (0.677), followed by retrieval-based approaches (0.661). However, MIL models showed larger generalization gaps, whereas DINOv2 lesion-crop retrieval remained stable (-0.014). Lesion-centered preprocessing was the dominant performance determinant within the retrieval paradigm. Embedding analysis revealed extensive class overlap, with 90.9% of fungal visits within the bacterial convex hull.
Conclusions:
Under leakage-aware, visit-level evaluation, all paradigms converged to modest discrimination, consistent with representational overlap within this single-center cohort. CNN-based MIL achieved higher AUROC but greater overfitting; retrieval-based models provided more stable generalization and superior probability calibration. Whether this ceiling is task-intrinsic or site-specific requires multicenter validation.
Translational Relevance:
Leakage-aware visit-level benchmarking supports realistic evaluation and calibrated decision support.