Related Experiment Videos
Diagnostic Performance of Deep Learning Models for Ocular Toxoplasmosis Using Fundus Photography: A Systematic Review
Seyed Mohammad Mousavi1, Masoud Rezayi1, Majid Fasihi Harandi1
1Research Center for Hydatid Disease in Iran, Institute of Infectious Diseases and Tropical Medicine, Department of Medical Parasitology, Afzalipour School of Medicine, Kerman University of Medical Sciences, Kerman, Iran.
Background:
Ocular toxoplasmosis (OT) is the leading infectious cause of posterior uveitis worldwide, yet diagnosis remains challenging without a universal gold standard. Although deep learning (DL) shows significant promise for automated image analysis, a rigorous synthesis of its diagnostic accuracy across heterogeneous architectures and datasets is lacking.
Methods:
Following PRISMA-DTA and CLAIM-AI guidelines, we searched major databases through January 2026 for studies evaluating DL models for OT detection using fundus photography. Methodological quality was appraised using QUADAS-2. A bivariate random-effects model (MetaDTA v2.1.3) was applied to pool sensitivity and specificity, generating hierarchical summary receiver operating characteristic (HSROC) curves to visualize discriminatory performance.
Results:
Sixteen primary studies (38 model configurations) met qualitative synthesis criteria; nine studies evaluating 24 unique architectures provided sufficient 2×2 contingency data for meta-analysis. Pooled sensitivity was 0.965 (95% CI: 0.944-0.978) and pooled specificity was 0.955 (95% CI: 0.929-0.972). The diagnostic odds ratio reached 594.1 (95% CI: 263.3-1340.4), with positive and negative likelihood ratios of 21.6 and 0.036. Substantial between-study heterogeneity was observed (I^2 = 87.9%). Critically, 81.2% of included studies relied exclusively on a single Paraguayan dataset (OTFID), raising concerns regarding phenotypic bias and limited global generalizability.
Conclusion:
DL models demonstrate excellent diagnostic accuracy for OT on fundus photography, comparable to established ophthalmic AI benchmarks. However, clinical translation is constrained by geographic data homogeneity and the absence of multicentric external validation. Future research must prioritize diverse international cohorts, explainable AI, and federated learning frameworks.