Related Experiment Video
Updated: Aug 28, 2025

09:09
Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
519
A novel silent speech recognition approach based on parallel inception convolutional neural network and Mel frequency
Jinghan Wu1,2, Yakun Zhang2,3, Liang Xie2,3
1Academy of Medical Engineering and Translational Medicine, Tianjin University, Tianjin, China.
Frontiers in Neurorobotics
|September 19, 2022
Summary
Silent speech recognition using surface electromyography (sEMG) signals achieves 90.76% accuracy. This novel approach, employing a Parallel Inception Convolutional Neural Network (PICNN), shows promise for practical applications in assistive technology.
Area of Science:
- Biomedical Engineering
- Machine Learning
- Human-Computer Interaction
Background:
- Silent speech recognition (SSR) aims to overcome limitations of traditional automatic speech recognition (ASR) where acoustic signals are unclear or absent.
- Current SSR methods require further development for real-world applicability.
- Surface electromyography (sEMG) offers a potential alternative signal source for silent speech detection.
Purpose of the Study:
- To develop a novel and effective silent speech recognition framework utilizing sEMG signals.
- To introduce and implement a new deep learning architecture, the Parallel Inception Convolutional Neural Network (PICNN), for sEMG-based SSR.
- To create a comprehensive dataset for silent speech recognition, focusing on daily life assistance demands.
Main Methods:
- A novel deep learning architecture, Parallel Inception Convolutional Neural Network (PICNN), was designed, processing six channels of sEMG data simultaneously.
- Mel Frequency Spectral Coefficients (MFSCs) were utilized for the first time to extract speech-related features from sEMG signals.
- A 100-class dataset was generated, encompassing demands relevant to elderly and disabled individuals.
Main Results:
- The proposed sEMG-based SSR system achieved a peak recognition accuracy of 90.76% across 28 subjects.
- The PICNN architecture outperformed existing state-of-the-art machine learning and deep learning algorithms in silent speech recognition.
- Subject-based transfer learning demonstrated effectiveness in enhancing cross-subject recognition capabilities.
Conclusions:
- The developed sEMG-based silent speech recognition system demonstrates high accuracy and robust performance.
- The novel PICNN architecture and MFSC feature extraction offer a significant advancement in the field of silent speech recognition.
- This technology holds considerable potential for practical applications, particularly in assistive technologies for individuals with communication impairments.
Related Concept Videos
Perceiving Loudness, Pitch, and Location
407
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
407
Parallel Processing
207
The brain processes sensory information rapidly due to parallel processing, which involves sending data across multiple neural pathways at the same time. This method allows the brain to manage various sensory qualities, such as shapes, colors, movements, and locations, all concurrently. For instance, when observing a forest landscape, the brain simultaneously processes the movement of leaves, the shapes of trees, the depth between them, and the various shades of green. This enables a quick and...
207
Linear Approximation in Frequency Domain
128
Linear systems are characterized by two main properties: superposition and homogeneity. Superposition allows the response to multiple inputs to be the sum of the responses to each individual input. Homogeneity ensures that scaling an input by a scalar results in the response being scaled by the same scalar.
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
128

