Related Experiment Video
Updated: Sep 3, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Sequence-to-Sequence Voice Reconstruction for Silent Speech in a Tonal Language.
Huiyan Li1, Haohong Lin1, You Wang1
1State Key Laboratory of Industrial Control Technology, Institute of Cyber Systems and Control, Zhejiang University, Hangzhou 310027, China.
This study introduces a novel sequence-to-sequence model for silent speech decoding (SSD) using surface electromyography (sEMG) in Mandarin Chinese. The approach successfully synthesizes voice from articulatory signals, achieving a low character error rate.
Area of Science:
- Neuroscience
- Speech Technology
- Brain-Computer Interfaces
Background:
- Silent speech decoding (SSD) utilizes articulatory neuromuscular activities for brain-computer interfaces (BCIs).
- Decoding surface electromyography (sEMG) for SSD is an active research area.
- Restoring silent speech in tonal languages like Mandarin Chinese presents significant challenges.
Purpose of the Study:
- To propose an optimized sequence-to-sequence (Seq2Seq) approach for synthesizing voice from sEMG-based silent speech.
- To address the difficulties in decoding silent speech for Mandarin Chinese.
- To improve the accuracy and naturalness of synthesized speech from articulatory signals.
Main Methods:
- Extraction of duration information from audio length to regulate sEMG-based silent speech.
- Implementation of a deep-learning encoder-decoder model.
- Utilization of a state-of-the-art vocoder for audio waveform generation.
Main Results:
- Successful decoding of silent speech in Mandarin Chinese across six speakers.
- Achievement of an average character error rate (CER) of 6.41%.
- Validation of the model's effectiveness through human evaluation.
Conclusions:
- The proposed optimized Seq2Seq model effectively synthesizes voice from sEMG signals for Mandarin Chinese silent speech.
- The method demonstrates significant potential for advancing BCIs and speech restoration technologies.
- This work contributes to overcoming challenges in decoding tonal languages for silent speech applications.
Related Concept Videos
Reconstruction of Signal using Interpolation
Larynx
Anatomy of the Larynx
The larynx consists of various components, including cartilage, muscles, and vocal cords. Its structure includes three large unpaired cartilages—the thyroid, cricoid, and epiglottis—and three smaller paired cartilages—the arytenoids,...
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Elaborative Rehearsals
The effectiveness of...
Components of Language
Chunking and Rehearsal in Sensory Memory

