Related Experiment Video
Updated: May 7, 2026

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
Parallel Multi-Attention and Gated Fusion for Visual Question Localized Answering in Surgical Scenes
Abstract:
Surgical Visual Question Localized Answering (Surgical-VQLA) is an emerging task that supports surgical education by generating accurate answers and localizing relevant anatomical regions based on visual content and textual queries. This task requires precise spatial reasoning and tight semantic alignment across modalities, which remain challenging for current models due to limited spatial sensitivity and insufficient semantic integration. Mitigating these limitations, we propose EndoVisLoc, a dedicated framework that enhances visual-textual interaction through structured attention and gated fusion. Specifically, we design a Parallel Multi Attention Module (PMAM) to capture different visual features, improving the perception of anatomical structures. We further develop a Dynamic Gated Fusion Module (DGFM) to adaptively inject semantic priors into visual features via gated control, facilitating robust cross-modal fusion. Finally, we introduce a Hierarchical Classifier Head (HCH) to refine the fused representations and jointly optimize answer prediction and spatial localization. Extensive experiments on the EndoVis-18-VQLA and EndoVis-17-VQLA datasets demonstrate the superior performance of EndoVisLoc, surpassing the state-of-the-art OTAS model by +5.72% ACC, +5.73% F-score, and +1.82% mIoU on EndoVis-18-VQLA and by +1.59% ACC, +2.48% F-score, and +0.33% mIoU on EndoVis-17-VQLA. These results confirm the consistent advantage of EndoVisLoc in both answer accuracy and precise anatomical localization.
