Temporal Pixel-Level Semantic Understanding Through the VSPW Dataset

Summary

This study introduces VSPW (Video Scene Parsing in the Wild), a large-scale dataset for video scene parsing. The proposed Temporal Attention Blending (TAB) Networks show superior performance for pixel-level semantic understanding in videos.