Related Experiment Video
Updated: Aug 5, 2026

04:48
Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
Scene Graph-Guided SegCaptioning Transformer With Fine-Grained Alignment for Controllable Video Segmentation and
Summary
This study introduces Controllable Video Segmentation and Captioning (SegCaptioning) for enhanced video understanding. The novel Scene Graph-guided Fine-grained SegCaptioning Transformer (SG-FSCFormer) model precisely interprets user intent for tailored multimodal outputs.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Natural Language Processing
Background:
- Multimodal large models advance video interpretation by generating correlated modalities.
- Existing methods lack user interaction, focusing on global video comprehension.
Purpose of the Study:
- Introduce Controllable Video Segmentation and Captioning (SegCaptioning) for user-guided video analysis.
- Enable simultaneous generation of masks and captions based on specific user prompts (e.g., bounding boxes).
Main Methods:
- Propose the Scene Graph-guided Fine-grained SegCaptioning Transformer (SG-FSCFormer) framework.
- Integrate a Prompt-guided Temporal Graph Former with an adaptive prompt adaptor to capture user intent.
- Utilize a Fine-grained Mask-linguistic Decoder and Multi-entity Contrastive loss for collaborative prediction and alignment.
Main Results:
- SG-FSCFormer demonstrates remarkable performance on benchmark datasets.
- The model effectively captures user intent and generates precise multimodal outputs.
- Achieved fine-grained alignment between generated masks and corresponding caption tokens.
Conclusions:
- SG-FSCFormer offers a novel approach to controllable video interpretation.
- The framework successfully addresses the limitations of existing methods by incorporating user interaction.
- Enables precise, user-specific video segmentation and captioning for enhanced comprehension.