Related Experiment Videos
Automatic generation of labanotation based on a hybrid transformer-LSTM network with multi-scale spatio-temporal
Huan Zhang1,2, Yan Li3,4, Kai Fan5
1School of Electrical and Information Technology, Yunnan Minzu University, Kunming, 650504, China.
Scientific Reports
|April 20, 2026
Summary
This study introduces a new model for automatically generating Labanotation from motion capture data, significantly improving the digital preservation of folk dance. The advanced MST-HTL model enhances accuracy and robustness in capturing complex dance movements.
Area of Science:
- Computer Science
- Digital Humanities
- Cultural Heritage Preservation
Background:
- Folk dance preservation requires accurate digital methods.
- Current Labanotation generation from motion capture has limitations in skeletal feature extraction and spatio-temporal dependency modeling.
Purpose of the Study:
- To develop an advanced model for accurate and robust automatic Labanotation generation.
- To enhance the digital preservation and dissemination of folk dance as intangible cultural heritage.
Main Methods:
- Proposed the Multi-Scale Spatio-Temporal Hybrid Transformer-LSTM (MST-HTL) model.
- Utilized a multi-scale spatio-temporal convolutional network for local skeletal feature extraction.
- Integrated a Transformer encoder and an LSTM decoder with enhanced attention for global spatio-temporal modeling.
Main Results:
- MST-HTL demonstrated superior performance over existing methods on benchmark datasets.
- Achieved a 0.92% improvement on LabanSeq16 and a 1.43% improvement on LabanSeq48.
- The model effectively captures fine-grained local features and global spatio-temporal dependencies.
Conclusions:
- The MST-HTL model offers a significant advancement in automatic Labanotation generation.
- Provides a strong foundation for digital preservation of dance heritage.
- Enhances the representation of complex dance movements for digital archiving.