Related Experiment Video
Updated: Jul 27, 2026

04:48
Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
2.8K
Parallel multi-head attention and term-weighted question embedding for medical visual question answering.
Sruthy Manmadhan1,2, Binsu C Kovoor1
1Division of Information Technology, Cochin University of Science and Technology, Kochi, Kerala 682022 India.
Summary
This study introduces a novel framework for medical visual question answering (Med-VQA) that bypasses the need for external data. The MaMVQA model achieves superior accuracy in answering clinical questions from medical images.
Area of Science:
- Artificial Intelligence
- Medical Imaging
- Natural Language Processing
Background:
- Medical Visual Question Answering (Med-VQA) models face challenges due to the unique nature of medical images and the scarcity of large-scale labeled datasets.
- Existing Med-VQA approaches often depend on transfer learning and cross-modal fusion, requiring external data for effective feature representation.
Purpose of the Study:
- To develop a novel parallel multi-head attention framework (MaMVQA) for Med-VQA that does not require external training data.
- To address the limitations of general domain Visual Question Answering (VQA) models in the medical context.
Main Methods:
- Image feature extraction is performed using an unsupervised Denoising Auto-Encoder (DAE).
- Language feature extraction utilizes term-weighted question embedding, enhanced by a supervised term-weighting (STW) scheme called qf-MI, based on mutual information (MI).
Main Results:
- The MaMVQA framework achieved state-of-the-art performance on the VQA-RAD benchmark without external data.
- The model demonstrated significantly improved accuracy for both close-ended (78.68%) and open-ended (55.31%) questions.
Conclusions:
- The proposed MaMVQA framework offers an effective solution for Med-VQA, overcoming data limitations.
- The study highlights the significance of individual components through extensive ablation studies, validating the framework's design.
Related Concept Videos
Depth Perception and Spatial Vision
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
Gestalt Principles of Perception
Gestalt principles provide a framework for understanding how humans perceive objects as unified wholes within their context. These principles are essential in explaining the cognitive processes that make sense of complex visual stimuli by organizing them into coherent groups. One fundamental principle is proximity, which posits that objects located close to each other are perceived as a collective group. For instance, when dots are positioned near one another, the visual system interprets them...
Parallel Processing
The brain processes sensory information rapidly due to parallel processing, which involves sending data across multiple neural pathways at the same time. This method allows the brain to manage various sensory qualities, such as shapes, colors, movements, and locations, all concurrently. For instance, when observing a forest landscape, the brain simultaneously processes the movement of leaves, the shapes of trees, the depth between them, and the various shades of green. This enables a quick and...
Prosopagnosia
Prosopagnosia, also known as face blindness, is the inability to recognize faces. In severe cases, individuals with prosopagnosia may not recognize close family members, including parents and spouses, by their faces. For instance, someone with prosopagnosia might walk past their child in a crowd, only realizing their mistake upon noticing their child's distinctive backpack or favorite jacket. Prosopagnosia specifically impairs facial recognition, while the recognition of other objects or...

