Scene Graph-Guided SegCaptioning Transformer With Fine-Grained Alignment for Controllable Video Segmentation and

Summary

This study introduces Controllable Video Segmentation and Captioning (SegCaptioning) for enhanced video understanding. The novel Scene Graph-guided Fine-grained SegCaptioning Transformer (SG-FSCFormer) model precisely interprets user intent for tailored multimodal outputs.

Related Concept Videos