用OphthalWeChat进行眼科视觉问题答案的大型多式模式的基准测试
Pusheng Xu1, Xia Gong1, Xiaolan Chen1
1School of Optometry, The Hong Kong Polytechnic University, Kowloon, Hong Kong, China.
Advances in ophthalmology practice and research
|January 29, 2026
概括
本研究介绍了OphthalWeChat,这是眼科的第一个双语视觉问题答案基准,用于评估视觉语言模型 (VLMs). 双子座2.0闪光灯实现了最高的精度,展示了眼部护理人工智能的进步.
科学领域:
- 眼科医生 眼科 眼科
- 人工智能的人工智能
- 医疗成像医学成像
背景情况:
- 视觉语言模型 (VLMs) 在医疗应用中表现有前途.
- 在眼科等专业领域评估VLM绩效需要量身定制的基准.
- 现有的基准可能无法捕捉现实世界临床数据的细微差别.
研究的目的:
- 为眼科开发一种双语多式视觉问答 (VQA) 基准.
- 为了促进视觉语言模型 (VLMs) 在眼科环境中的评估.
- 支持创建用于眼科护理的先进AI系统.
主要方法:
- 从微信官方帐户 (2016-2024) 收集眼科图像帖子和标题.
- 使用GPT-4o-mini.生成双语 (中英) 问答对使用GPT-4o-mini.
- 分类QA对成二进制,单一选择和开放式子集.
- 评估了六个VLM (GPT-4o,双子座2.0闪光灯等) 使用精度和基于语言的指标.
主要成果:
- 眼科微信基准包括3469个图像和30120个QA对,跨越9个子专业.
- 双子 2.0 闪光实现了最高的整体精度 (0.555),超过了其他 VLM.
- 性能因问题类型,语言,子专业和成像模式而异;损伤/诊断错误最常见.
结论:
- 眼科微信是第一个双语的眼科VQA基准,使用真实世界的数据.
- 该基准允许对眼科护理中的VLM进行定量评估.
- 该资源有助于开发专业和准确的眼科人工智能系统.
相关概念视频
Visual System
1.8K
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
Once through the pupil, the light passes through the lens, a...
1.8K
Visual Agnosia
1.1K
Visual agnosia is a condition characterized by the inability to recognize visually presented objects despite having normal vision. For instance, a person with visual agnosia can describe the shape and color of an object but cannot identify or name it. This impairment does not affect their visual field, acuity, color vision, brightness discrimination, language, or memory. An example of this condition in a social setting is someone at a dinner party asking for "that silver thing with a round...
1.1K
Molecular Models
43.7K
Physical models representing molecular architectures of chemical compounds play essential roles in understanding chemistry. The use of molecular models makes it easier to visualize the structures and shapes of atoms and molecules.
43.7K
Photoreceptors and Visual Pathways
9.2K
At the molecular level, visual signals trigger transformations in photopigment molecules, resulting in changes in the photoreceptor cell's membrane potential. The photon's energy level is denoted by its wavelength, with each specific wavelength of visible light associated with a distinct color. The spectral range of visible light, classified as electromagnetic radiation, spans from 380 to 720 nm. Electromagnetic radiation wavelengths exceeding 720 nm fall under the infrared category,...
9.2K
The Bohr Model
80.8K
Following the work of Ernest Rutherford and his colleagues in the early twentieth century, the picture of atoms consisting of tiny dense nuclei surrounded by lighter and even tinier electrons continually moving about the nucleus was well established. This picture was called the planetary model since it pictured the atom as a miniature “solar system” with the electrons orbiting the nucleus like planets orbiting the sun. The simplest atom is hydrogen, consisting of a single proton as the...
80.8K
Stereotype Content Model
15.4K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
15.4K


