现成的视觉大语言模型可以从视网膜照片中检测和诊断眼睛疾病吗?
Sahana Srinivasan1,2,3, Hongwei Ji4, David Ziyou Chen1,5
1Centre for Innovation and Precision Eye Health, Yong Loo Lin School of Medicine, National University of Singapore and National University Health System, Singapore.
BMJ open ophthalmology
|April 7, 2025
概括
视觉大语言模型 (VLLMs) 在检测视网膜图像中的眼睛异常方面表现出高准确性. 然而,它们的诊断能力目前还不足以用于临床使用,这凸显了对专门模型和人类监督的需求.
科学领域:
- 眼科医生 眼科 眼科
- 人工智能的人工智能
- 医疗成像医学成像
背景情况:
- 生成型人工智能已经导致视觉大语言模型 (VLLMs) 的发展.
- 评估这些VLLM在眼科中的临床实用性至关重要.
研究的目的:
- 评估OpenAI的GPT-4V和Google Gemini在检测和诊断视网膜图像中的眼睛疾病方面的表现.
- 为了比较不同VLLM配置的准确性和诊断质量.
主要方法:
- 利用了来自新加坡眼病流行病学 (SEED) 研究的44张视网膜图像,包括健康对照和六种特定的眼病.
- 提示GPT-4V (默认和数据分析模式) 和谷歌双子来识别视网膜异常并提供诊断.
- 眼科医生评估了VLLM输出的准确性和诊断描述质量.
主要成果:
- GPT-4V默认模式实现了最高的检测率 (97.1%),明显超过了其数据分析模式 (61.8%) 和谷歌双子 (41.2%).
- 尽管检测率很高,但在所有VLLM中,诊断描述的质量普遍很差,只有4.8%-28.6%被评为好.
- GPT-4V默认模式显示出对异常检测的高灵敏度,但缺乏诊断准确性.
结论:
- 虽然VLLMs在检测眼部异常方面具有潜力,但它们目前的诊断能力不足以用于临床应用.
- 域特定的VLLM是必要的,以提高诊断准确度.
- 在临床眼科中,人类专家的监督对于准确的诊断和患者护理至关重要.
相关概念视频
The Retina
66.4K
The retina is a layer of nervous tissue at the back of the eye that transduces light into neural signals. This process, called phototransduction, is carried out by rod and cone photoreceptor cells in the back of the retina.
66.4K
Vision
52.4K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
52.4K


