TSMS-SAM2:手術シナリオにおけるプロンプト可能なビデオオブジェクトセグメンテーションおよびトラッキングのためのマルチスケール時間サンプリング拡張とメモリ分割プルーニング
Guoping Xu1, Hua-Chieh Shao1, You Zhang1
1The Medical Artificial Intelligence and Automation (MAIA) Laboratory, Department of Radiation Oncology, University of Texas Southwestern Medical Center, Dallas, TX 75390, USA.
Abstract:
Promptable video object segmentation and tracking (VOST) has seen significant advances with the emergence of foundation models like Segment Anything Model 2 (SAM2); however, their application in surgical video analysis remains challenging due to complex motion dynamics and the redundancy of memory that impedes effective learning. In this work, we propose TSMS-SAM2, a novel framework that enhances promptable VOST in surgical videos by addressing challenges of rapid object motion and memory redundancy in SAM2. TSMS-SAM2 introduces two key strategies: multi-temporal-scale video sampling augmentation to improve robustness against motion variability, and a memory splitting and pruning mechanism that organizes and filters past frame features for more efficient and accurate segmentation. Evaluated on EndoVis2017 and EndoVis2018 datasets, TSMS-SAM2 achieved the highest mean (± s.d.) Dice scores of 95.24±0.96% and 86.73±15.46%, respectively, outperforming prior SAM-based and task-specific methods. Extensive ablation studies confirm the effectiveness of multiscale temporal augmentation and memory splitting, highlighting the framework's potential for robust, efficient segmentation in complex surgical scenarios. Our source code will be made available at https://github.com/apple1986/TSMS-SAM2.
関連する概念動画
pH Scale
System of Memory
Working Memory
¹H NMR: Complex Splitting
Splitting diagrams or splitting tree diagrams are routinely used to depict such complex couplings. While drawing splitting diagrams, the splitting with the larger coupling constant is usually applied...
Potential Due to a Polarized Object
Potential Due to a Magnetized Object
The vector...


