Related Experiment Videos
State-change learning for prediction of future events in endoscopic videos
Saurav Sharma1, Chinedu Innocent Nwoye1, Didier Mutter2
1University of Strasbourg, CNRS, INSERM, ICube, UMR7357, France; IHU Strasbourg, France.
Abstract:
Surgical future prediction, driven by real-time AI analysis of surgical video, is critical for operating room safety and efficiency. Future prediction could provide actionable insights into upcoming events, their timing, and associated risks-enabling better resource allocation, timely instrument readiness, and early warnings for emergent complications (e.g., bleeding, bile duct injury). Despite this need, current surgical AI research focuses on understanding what is happening rather than predicting future events. Existing methods target specific tasks such as phase or instrument anticipation in isolation, lacking unified approaches that span both short-term (action triplets, surgical events) and long-term horizons (remaining surgery duration, phase/step transitions). These methods rely on coarse-grained supervision at the phase or instrument level, while fine-grained surgical action triplets and steps remain underexplored despite their potential to capture nuanced temporal dynamics. We address these limitations by reframing surgical future prediction as state-change learning. Rather than forecasting raw observations directly, our approach classifies state transitions between current and future timesteps, building transition-aware representations that improve generalization across tasks and procedures. In this work, we introduce SurgFUTR, implementing this paradigm through a teacher-student architecture. Video clips are compressed into state representations via Sinkhorn-Knopp clustering; the teacher network learns from both current and future clips, while the student network predicts future states from current observations alone, guided by our Action Dynamics (ActDyn) module that models state transition patterns. For comprehensive evaluation, we establish SFPBench, spanning five prediction tasks across different temporal horizons: short-term anticipation (cystic-structure triplets, surgical events) and long-term forecasting (remaining surgery duration, phase/step transitions). Across four datasets spanning three laparoscopic procedures, SurgFUTR provides generally favorable and often competitive performance relative to strong baselines, with the clearest gains on several long-horizon and fine-grained anticipation tasks. Cross-procedure transfer from cholecystectomy to gastric bypass further shows encouraging, though task and center-dependent, generalization. Code will be made available at https://github.com/CAMMA-public/surgfutr.