Related Experiment Video
Updated: Jul 16, 2026

08:15
Capturing Dynamic Finger Gesturing with High-resolution Surface Electromyography and Computer Vision
Published on: March 28, 2025
Text-to-Korean Sign Language Pose Sequence Generation Using Non-Manual Signal Conditioning and Multi-Scale Temporal
1Department of Smart ICT Convergence Engineering, Seoul National University of Science and Technology, 232 Gongneung-ro, Nowon-gu, Seoul 01811, Republic of Korea.
Sensors (Basel, Switzerland)
|July 15, 2026
Summary
This study introduces a new model for generating Korean Sign Language (KSL) poses from text, improving accuracy and naturalness by considering non-manual signals and temporal coherence. The advanced text-to-pose model enhances information accessibility for the deaf and hard-of-hearing community.
Area of Science:
- Computer Science
- Artificial Intelligence
- Linguistics
Background:
- Automatic sign language generation aids information accessibility for deaf and hard-of-hearing individuals.
- Generating sign language poses from text is complex due to manual and non-manual signals, and temporal coherence requirements.
- Existing methods struggle with the non-linear relationship between text length and pose sequence duration.
Purpose of the Study:
- To propose a novel text-to-Korean Sign Language (KSL) pose generation model.
- To address challenges in sign language generation, including non-manual signals and temporal dynamics.
- To improve the accuracy and naturalness of avatar-based sign language expression.
Main Methods:
- Developed a framework integrating text encoding, pose decoding, non-manual signal conditioning, and multi-scale temporal refinement.
- The model generates normalized 58-joint KSL keypoint sequences from morpheme-level text.
- Joint optimization included pose reconstruction, motion continuity, bone consistency, precision, non-manual signal prediction, and length consistency.
Main Results:
- The proposed model significantly outperformed text-only and Transformer-based baselines in KSL pose generation.
- Key metrics improved, including reduced Mean Per Joint Position Error (MPJPE) and Pose Mean Absolute Error (MAE).
- Non-manual signal prediction (F1 score) saw a substantial increase, indicating better integration of facial expressions and body posture.
Conclusions:
- Text-based KSL pose generation necessitates a holistic approach, integrating non-manual expressions, length consistency, and long-term temporal structure.
- The model demonstrates a significant advancement over frame-wise keypoint prediction methods.
- Further research is needed to validate linguistic meaningfulness and real-world accessibility beyond coordinate-level improvements.
