Related Experiment Videos
Foundation Models in Ophthalmic Artificial Intelligence: Current Status and Future Directions
Abstract:
Recent advances in ophthalmic foundation models have accelerated the application of artificial intelligence in ophthalmology, improving disease detection, progression assessment, and treatment evaluation. These advances are supported by the increasing availability of multimodal ophthalmic imaging data, including optical coherence tomography, color fundus photography, and slit lamp imaging, together with progress in self-supervised learning, vision-language modeling, and agent-based systems. This review summarizes publicly available multimodal ophthalmic imaging datasets and examines their foundational role in large-scale pretraining and cross-modal representation learning. We organize ophthalmic foundation models into vision foundation models, vision-language foundation models, and Large Language Model-based agentic systems, and discuss their core architectures, learning strategies, and representative applications. By synthesizing evidence across structural, vascular, and anterior-segment modalities, we illustrate how ophthalmic artificial intelligence is evolving from task-specific image recognition toward semantically grounded, clinically oriented cognition and interactive reasoning. We further highlight key challenges that hinder clinical translation, including data imbalance across modalities, imperfect image-text semantic alignment, limited robustness and generalizability, hallucinations in generative agents, and barriers to clinical workflow integration. Finally, we outline future directions, emphasizing standardized multimodal data resources, improved interpretability and factual consistency, rigorous multi-center validation, and cross-disciplinary collaboration to support reliable clinical decision-making.