Related Experiment Video
Updated: Jul 15, 2026

Haptic/Graphic Rehabilitation: Integrating a Robot into a Virtual Environment Library and Applying it to Stroke Therapy
Published on: August 8, 2011
Dual modality prompt learning for visual question-grounded answering in robotic surgery
Yue Zhang1, Wanshu Fan2, Peixi Peng1
1National and Local Joint Engineering Laboratory of Computer Aided Design, School of Software Engineering, Dalian University, Dalian, 116622, Liaoning, China.
This study introduces a grounded visual question answering (VQA) model for robotic surgery that precisely locates relevant image regions while answering questions. The novel dual-modality prompt model enhances multimodal interactions for improved accuracy in surgical contexts.
Area of Science:
- Robotic Surgery
- Computer Vision
- Artificial Intelligence
Background:
- Advancements in robotic surgery have spurred progress in visual question answering (VQA).
- Current VQA systems lack the ability to localize relevant image regions, limiting interpretability and exploration in surgical contexts.
Purpose of the Study:
- To propose a grounded VQA model for robotic surgery that can localize specific regions during answer prediction.
- To enhance multimodal information interactions for precise localization and accurate answer inference.
Main Methods:
- Developed a dual-modality prompt model inspired by prompt learning in language models.
- Introduced visual and textual complementary prompters for integrating visual and textual information.
- Employed a multiple iterative fusion strategy for comprehensive answer reasoning.
Main Results:
- The proposed grounded VQA model demonstrated superior performance compared to existing methods.
- The model effectively integrates visual and textual prompts for accurate localization and answer generation.
- Experiments were validated on the EndoVis-18 and EndoVis-17 datasets.
Conclusions:
- The grounded VQA model significantly improves localization and answer prediction in robotic surgery.
- The dual-modality prompt approach enhances multimodal interactions for VQA tasks.
- This model offers a more interpretable and exploratory VQA solution for surgical applications.
More Related Videos
05:28Author Spotlight: Enhancing Upper Limb Rehabilitation in Stroke Patients Through Advanced Robotic and Neuromodulation Technologies
Published on: October 11, 2024
07:46Author Spotlight: Revolutionizing Remote Surgery with Augmented Reality and Robotics for Enhanced Precision and Accessibility
Published on: August 9, 2024