Related Experiment Video
Updated: Sep 11, 2025

05:48
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
1.6K
Scaling up Multimodal Pre-Training for Sign Language Understanding
IEEE Transactions on Pattern Analysis and Machine Intelligence
|August 14, 2025
Summary
This study introduces a multimodal sign language pre-training (SLP) framework using a large dataset (SL-1.5M) to improve sign language understanding (SLU) models. The new method enhances model generalization by integrating visual and textual cues for better sign language video representation.
Area of Science:
- Computer Science
- Artificial Intelligence
- Natural Language Processing
Background:
- Sign language pre-training (SLP) enhances sign language understanding (SLU) but faces limitations in model generalization and neglecting textual cues.
- Existing methods often use task-specific pre-training on small datasets or focus only on visual information, reducing model representational capacity.
Purpose of the Study:
- To develop a multimodal SLP framework that leverages visual context and vision-language consistency for improved sign language video representation.
- To address data scarcity by curating a large-scale text-labeled sign pose dataset (SL-1.5M).
Main Methods:
- Curated a large-scale text-labeled sign pose dataset (SL-1.5M) from diverse sources.
- Proposed a pre-training framework integrating sign-text contrastive learning and masked pose modeling.
- Concurrently modeled manual and non-manual sign language information for holistic visual content representation.
Main Results:
- The framework effectively captures contextual cues in sign pose sequences and aligns semantic text features.
- Achieved new state-of-the-art performance on diverse SLU tasks, demonstrating superior generalization and effectiveness.
- Validated the framework's ability to enhance the representative capability of sign language videos.
Conclusions:
- The proposed multimodal SLP framework significantly improves sign language understanding by integrating visual and textual information.
- The approach overcomes limitations of existing methods, offering better generalization and representational capacity for sign language models.
- This work sets a new benchmark for sign language understanding through advanced pre-training techniques.

