Related Experiment Video
Updated: Jul 16, 2026

Capturing Dynamic Finger Gesturing with High-resolution Surface Electromyography and Computer Vision
Published on: March 28, 2025
Text-to-Korean Sign Language Pose Sequence Generation Using Non-Manual Signal Conditioning and Multi-Scale Temporal
1Department of Smart ICT Convergence Engineering, Seoul National University of Science and Technology, 232 Gongneung-ro, Nowon-gu, Seoul 01811, Republic of Korea.
Abstract:
Automatic sign language generation has the potential to support information accessibility for deaf and hard-of-hearing individuals. Generating sign language pose sequences from natural language text can serve as an intermediate representation for avatar-based sign language expression and sign language video synthesis. However, text-to-sign pose generation is challenging because sign language conveys meaning through both manual movements and non-manual signals, while requiring temporally coherent motion over local and sentence-level contexts. In addition, text length does not directly correspond to the number of pose frames required for sign language expression. To address these issues, this study proposes a text-to-Korean Sign Language (KSL) pose generation model based on non-manual signal conditioning and multi-scale temporal refinement. The proposed framework integrates a text encoder, pose decoder, non-manual signal conditioning, multi-scale temporal refinement, and length prediction/blending. The model generates normalized 58-joint KSL keypoint sequences from morpheme-level text inputs and jointly optimizes pose reconstruction, motion continuity, bone consistency, PCK-aware precision, non-manual signal prediction, and length consistency. Experimental results on a KSL text-pose dataset show that the proposed model outperforms text-only and Transformer-based baselines. Compared with the Transformer text-to-pose baseline, the proposed model reduced MPJPE from 0.408236 to 0.316366 and Pose MAE from 0.165473 to 0.128570. It also improved PCK@0.05 from 0.136090 to 0.163928 and reduced the length relative error from 0.221455 to 0.127152. In particular, the best-threshold non-manual F1 substantially increased from 0.010859 to 0.494566. These results suggest that text-based KSL pose generation should jointly consider non-manual expressions, length consistency, and long-term temporal motion structure rather than relying only on frame-wise keypoint prediction. However, the reported improvements should be interpreted as coordinate- and label-level evidence, not as a complete validation of linguistic meaningfulness or real-world accessibility.
