Referring Segmentation in Images and Videos With Cross-Modal Self-Attention Network

Related Concept Videos