Surgical Video Understanding with Alignment-Preserving Temporal Adaptation and Action Triplet Text Alignment

Taiyo Ikeido1, Ren Togo2, Takahiro Ogawa2

  • 1Graduate School of Information Science and Technology, Hokkaido University, Kita 14, Nishi 9, Kita-ku, Sapporo 060-0814, Hokkaido, Japan.

Summary

This study introduces an efficient framework for surgical video analysis using a pretrained vision-language model and a temporal adapter. It enhances surgical phase recognition and enables few-shot activity recognition with minimal annotations.

Related Concept Videos