报告生成系统用于使用视觉语言模型的裂灯图像解释
Xin Ye1, Yingjiao Shen1, Qian Chen2
1Department of Ophthalmology, Zhejiang Provincial People's Hospital (Affiliated People's Hospital, Hangzhou Medical College), Hangzhou, Zhejiang, China.
Ophthalmology and therapy
|March 14, 2026
概括
这项研究开发了使用视觉语言模型 (VLMs) 来生成医疗报告的裂纹灯 (SL) 图像的解释管道. 开发的LLaVA和Qwen2.5-VL框架在协助眼科医生进行眼科图像解释方面显示出有希望的结果.
科学领域:
- 眼科医生 眼科 眼科
- 人工智能的人工智能
- 医疗成像医学成像
背景情况:
- 裂纹灯 (SL) 成像对于诊断眼科疾病至关重要.
- 自动解释SL图像可以提高诊断效率.
- 视觉语言模型 (VLMs) 为图像解释和报告生成提供了潜力.
研究的目的:
- 开发和评估使用VLMs生成SL图像报告的解释管道.
- 微调和评估眼科图像分析的LLaVA和Qwen2.5-VL框架的性能.
- 探索VLM在协助眼科医生和增强患者护理方面的潜力.
主要方法:
- 使用引导语言图像预训练 (BLIP) 开发了一个图像-文本对齐模块.
- 精细调整的LLaVA和Qwen2.5-VL框架在一个SL图像和湖医院的医疗报告数据集上.
- 通过Bijie医院的外部数据集验证了框架.
主要成果:
- 改进后的LLaVA和Qwen2.5-VL模型在报告生成方面取得了很高的表现,LLaVA在正确性,完整性和满意度方面的得分略有提高.
- 两种模型都表现出强大的疾病分类准确度 (0.87-0.88),特别是对于常见的疾病,如玻璃眼和结膜炎.
- 眼科医生之间观察者间的一致性很大 (k分数为0.714-0.777).
结论:
- 开发的框架有效地生成SL图像的报告,增强眼科图像解释.
- 在临床实践中,VLM显示出有很大的潜力来帮助眼科医生.
- 这项技术可以提高诊断准确度,并简化眼科疾病的报告.
相关概念视频
Vision
61.2K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
61.2K
Confocal Fluorescence Microscopy
21.7K
Confocal microscopy is an advanced microscopic technique. The prime advantage of the confocal microscope over other microscopy techniques is its ability to block the out-of-focus light from the illuminated samples using pinholes. It is widely used with fluorescence optics to obtain high-resolution, sharp contrast images. Unlike optical microscopes, confocal microscopes use a focused beam of light laser to scan the entire sample surface at different z-planes. These microscopes are, therefore,...
21.7K


