Related Experiment Video
Updated: Jul 29, 2025

07:46
Author Spotlight: Revolutionizing Remote Surgery with Augmented Reality and Robotics for Enhanced Precision and Accessibility
Published on: August 9, 2024
787
LAST: LAtent Space-Constrained Transformers for Automatic Surgical Phase Recognition and Tool Presence Detection
IEEE Transactions on Medical Imaging
|May 25, 2023
Summary
This study introduces LAST, a novel multi-task learning method for surgical phase recognition and tool detection. LAST improves accuracy by leveraging video-level semantic information, outperforming existing methods on public datasets.
Area of Science:
- Medical image analysis
- Artificial intelligence in surgery
- Computer vision
Background:
- Automatic surgical phase recognition and tool presence detection are crucial for context-aware surgical systems.
- Existing methods often use frame-level loss functions, failing to capture the full semantic structure of surgeries.
- This leads to suboptimal performance in surgical video analysis.
Purpose of the Study:
- To propose a novel multi-task learning framework called LAST (LAtent Space-constrained Transformers).
- To enhance surgical phase recognition and tool presence detection by leveraging video-level semantic information.
- To improve the accuracy and robustness of surgical context-aware systems.
Main Methods:
- Developed a two-branch transformer architecture incorporating multi-task learning.
- Introduced a novel method to leverage video-level semantic information using a transformer variational autoencoder (VAE).
- Ensured predictions align with learned statistical distributions and lie on an extracted low-dimensional data manifold, making the model structure-aware.
Main Results:
- Achieved superior performance on the Cholec80 and M2cai16 datasets compared to state-of-the-art methods.
- On Cholec80, reported average accuracies of 93.12% for phase recognition and 95.15% mAP for tool presence detection.
- Demonstrated significant improvements in precision, recall, and Jaccard index for phase recognition.
Conclusions:
- The proposed LAST method effectively integrates video-level semantic information for surgical context-aware tasks.
- LAST demonstrates state-of-the-art performance in both surgical phase recognition and tool presence detection.
- The structure-aware approach significantly enhances the capabilities of automated surgical analysis systems.

