Related Experiment Video
Updated: Jan 11, 2026

09:27
Using Eye Movements Recorded in the Visual World Paradigm to Explore the Online Processing of Spoken Language
Published on: October 13, 2018
10.6K
Probing the Visualization Literacy of Vision Language Models: The Good, the Bad, and the Ugly
IEEE Transactions on Visualization and Computer Graphics
|November 19, 2025
Summary
Vision Language Models (VLMs) show strong chart understanding. Our study uses attention-guided class activation maps (AG-CAM) to reveal VLM reasoning, finding ChartGemma excels among open-source models.
Area of Science:
- Artificial Intelligence
- Computer Vision
- Natural Language Processing
Background:
- Vision Language Models (VLMs) show potential in chart comprehension.
- Previous research primarily assessed VLM response accuracy, neglecting internal reasoning processes.
- Understanding VLM visualization literacy requires exploring their decision-making mechanisms.
Purpose of the Study:
- To adapt and apply attention-guided class activation maps (AG-CAM) for visualizing VLM internal reasoning in chart comprehension.
- To compare the performance and internal reasoning of open-source and closed-source VLMs on chart-based tasks.
- To investigate VLM spatial and semantic reasoning capabilities during chart question-answering.
Main Methods:
- Adapted attention-guided class activation maps (AG-CAM) for early fusion VLM architectures.
- Evaluated four open-source (ChartGemma, Janus 1B/7B, LLaVA) and two closed-source (GPT-4o, Gemini) VLMs.
- Analyzed AG-CAM results to understand feature importance and model reasoning.
Main Results:
- ChartGemma, a 3B parameter VLM, outperformed other open-source models and matched closed-source models in performance.
- VLMs demonstrated spatial reasoning by localizing chart features and semantic reasoning by linking visual elements to data.
- AG-CAM visualization revealed insights into VLM decision-making processes for chart QA.
Conclusions:
- The adapted AG-CAM method provides a novel approach for analyzing VLM internal reasoning in chart comprehension.
- Open-source VLMs, particularly ChartGemma, show competitive performance and interpretable reasoning capabilities.
- This work advances transparent and reproducible research in AI visualization literacy and chart QA.
More Related Videos
Related Concept Videos
Vision
59.2K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
59.2K
Depth Perception and Spatial Vision
1.8K
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
1.8K
Visual Agnosia
912
Visual agnosia is a condition characterized by the inability to recognize visually presented objects despite having normal vision. For instance, a person with visual agnosia can describe the shape and color of an object but cannot identify or name it. This impairment does not affect their visual field, acuity, color vision, brightness discrimination, language, or memory. An example of this condition in a social setting is someone at a dinner party asking for "that silver thing with a round...
912
Visual System
1.6K
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
Once through the pupil, the light passes through the lens, a...
1.6K

