Related Experiment Video
Updated: Jan 9, 2026

04:48
Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
Published on: July 5, 2024
725
VMSIS: A Pre-trained Vision Transformer with Mamba Decoder for Surgical Instrument Segmentation.
Summary
We developed VMSIS, a novel hybrid model for robot-assisted surgery instrument segmentation. It combines DINOv2 and Mamba to accurately identify surgical tools in videos while maintaining temporal consistency.
Area of Science:
- Robotics
- Computer Vision
- Medical Imaging
Background:
- Accurate surgical instrument segmentation is crucial for robot-assisted surgery.
- Existing methods may struggle with temporal consistency and efficiency.
Purpose of the Study:
- To introduce VMSIS, a hybrid architecture for enhanced surgical instrument segmentation.
- To leverage self-supervised learning and efficient sequence modeling for improved performance.
Main Methods:
- A hybrid architecture combining DINOv2 (frozen backbone) and a Mamba-based decoder.
- Training on over 900,000 frames of RGB surgical videos.
- Processing 10 consecutive frames to capture temporal dependencies.
Main Results:
- Achieved accurate surgical instrument segmentation with temporal consistency.
- Demonstrated effectiveness across 4 reorganized public datasets.
- Obtained competitive results with fewer trainable parameters than traditional methods.
Conclusions:
- VMSIS offers an effective approach for surgical instrument segmentation in robot-assisted surgery.
- The hybrid DINOv2-Mamba architecture successfully captures visual and temporal information.
- The model shows promise for improving surgical precision and safety.

