Related Experiment Video
Updated: Jan 30, 2026

09:56
In Vivo Multimodal Imaging and Analysis of Mouse Laser-Induced Choroidal Neovascularization Model
Published on: January 21, 2018
9.9K
Benchmarking large multimodal models for ophthalmic visual question answering with OphthalWeChat
Pusheng Xu1, Xia Gong1, Xiaolan Chen1
1School of Optometry, The Hong Kong Polytechnic University, Kowloon, Hong Kong, China.
Advances in Ophthalmology Practice and Research
|January 29, 2026
Summary
This study introduces OphthalWeChat, the first bilingual visual question answering benchmark for ophthalmology, to evaluate vision-language models (VLMs). Gemini 2.0 Flash achieved the highest accuracy, demonstrating advancements in AI for eye care.
Area of Science:
- Ophthalmology
- Artificial Intelligence
- Medical Imaging
Background:
- Vision-language models (VLMs) show promise in medical applications.
- Evaluating VLM performance in specialized fields like ophthalmology requires tailored benchmarks.
- Existing benchmarks may not capture the nuances of real-world clinical data.
Purpose of the Study:
- To develop a bilingual multimodal visual question answering (VQA) benchmark for ophthalmology.
- To facilitate the evaluation of Vision-language Models (VLMs) in ophthalmic contexts.
- To support the creation of advanced AI systems for eye care.
Main Methods:
- Collected ophthalmic image posts and captions from WeChat Official Accounts (2016-2024).
- Generated bilingual (Chinese/English) question-answer pairs using GPT-4o-mini.
- Categorized QA pairs into binary, single-choice, and open-ended subsets.
- Evaluated six VLMs (GPT-4o, Gemini 2.0 Flash, etc.) using accuracy and language-based metrics.
Main Results:
- The OphthalWeChat benchmark comprises 3469 images and 30120 QA pairs across 9 subspecialties.
- Gemini 2.0 Flash achieved the highest overall accuracy (0.555), outperforming other VLMs.
- Performance varied by question type, language, subspecialty, and imaging modality; lesion/diagnosis errors were most frequent.
Conclusions:
- OphthalWeChat is the first bilingual VQA benchmark for ophthalmology, using real-world data.
- The benchmark enables quantitative evaluation of VLMs in eye care.
- This resource aids in developing specialized and accurate AI systems for ophthalmology.
Related Concept Videos
Visual System
1.8K
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
Once through the pupil, the light passes through the lens, a...
1.8K
Visual Agnosia
1.1K
Visual agnosia is a condition characterized by the inability to recognize visually presented objects despite having normal vision. For instance, a person with visual agnosia can describe the shape and color of an object but cannot identify or name it. This impairment does not affect their visual field, acuity, color vision, brightness discrimination, language, or memory. An example of this condition in a social setting is someone at a dinner party asking for "that silver thing with a round...
1.1K
Molecular Models
43.7K
Physical models representing molecular architectures of chemical compounds play essential roles in understanding chemistry. The use of molecular models makes it easier to visualize the structures and shapes of atoms and molecules.
43.7K
Photoreceptors and Visual Pathways
9.2K
At the molecular level, visual signals trigger transformations in photopigment molecules, resulting in changes in the photoreceptor cell's membrane potential. The photon's energy level is denoted by its wavelength, with each specific wavelength of visible light associated with a distinct color. The spectral range of visible light, classified as electromagnetic radiation, spans from 380 to 720 nm. Electromagnetic radiation wavelengths exceeding 720 nm fall under the infrared category,...
9.2K
The Bohr Model
80.8K
Following the work of Ernest Rutherford and his colleagues in the early twentieth century, the picture of atoms consisting of tiny dense nuclei surrounded by lighter and even tinier electrons continually moving about the nucleus was well established. This picture was called the planetary model since it pictured the atom as a miniature “solar system” with the electrons orbiting the nucleus like planets orbiting the sun. The simplest atom is hydrogen, consisting of a single proton as the...
80.8K
Stereotype Content Model
15.4K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
15.4K

