Scaling up Multimodal Pre-Training for Sign Language Understanding

Summary

This study introduces a multimodal sign language pre-training (SLP) framework using a large dataset (SL-1.5M) to improve sign language understanding (SLU) models. The new method enhances model generalization by integrating visual and textual cues for better sign language video representation.