Related Experiment Video
Updated: Apr 13, 2026

07:35
A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
Published on: October 13, 2023
1.5K
Enhancing Visual Reasoning With LLM-Powered Knowledge Graphs for Visual Question Localized-Answering in Robotic
IEEE Journal of Biomedical and Health Informatics
|March 3, 2025
Summary
This study introduces a new framework to improve surgical visual question answering by using Large Language Model-powered knowledge graphs for better instrument recognition. This enhances medical training by providing clearer visual reasoning for surgical procedures.
Area of Science:
- Medical Imaging
- Artificial Intelligence in Surgery
- Computer Vision
Background:
- Expert surgeons face heavy workloads, limiting timely responses to trainees.
- Surgical Visual Question Localized-Answering (Surgical-VQLA) is crucial for medical education.
- Current Surgical-VQLA models struggle with accurate surgical instrument identification due to limited visual reasoning.
Purpose of the Study:
- To enhance visual reasoning capabilities in Surgical-VQLA.
- To improve the accurate identification and understanding of surgical instruments within surgical scenes.
- To develop a framework that leverages external knowledge to augment visual understanding in surgical contexts.
Main Methods:
- Proposed Enhancing Visual Reasoning with LLM-Powered Knowledge Graphs (EnVR-LPKG) framework.
- Developed a Fine-grained Knowledge Extractor (FKE) for targeted knowledge graph information extraction.
- Introduced a Multi-attention-based Surgical Instrument Enhancer (MSIE) module for fusing visual and knowledge graph features.
Main Results:
- The EnVR-LPKG framework significantly improved the accuracy of surgical instrument identification.
- The MSIE module effectively fused visual and textual knowledge graph features, enhancing model understanding.
- Experimental results on EndoVis-17-VQLA and EndoVis-18-VQLA datasets showed superior performance compared to state-of-the-art methods.
Conclusions:
- The proposed EnVR-LPKG framework effectively addresses limitations in current Surgical-VQLA models.
- LLM-powered knowledge graphs enhance visual reasoning for surgical instrument recognition.
- This approach offers a promising direction for improving surgical training and education through AI.

