Related Experiment Video
Updated: Jul 12, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Speech extraction from vibration signals based on deep learning
Li Wang1,2,3, Weiguang Zheng1,4, Shande Li1,3
1State Key Laboratory of Digital Manufacturing Equipment and Technology, School of Mechanical Science and Engineering, Huazhong University of Science and Technology, Wuhan, China.
This study introduces a deep learning method for speech extraction from vibration signals, overcoming traditional limitations. Fully connected networks demonstrate superior performance and robustness in speech recognition from vibroacoustic data.
Area of Science:
- Acoustics and Signal Processing
- Machine Learning and Artificial Intelligence
- Structural Dynamics and System Identification
Background:
- Traditional speech extraction from vibration signals is sensitive to deviations in model parameters, noise, and boundary conditions.
- Existing methods face challenges in accurately identifying speech characteristics from complex vibroacoustic systems.
- Developing robust methods for speech extraction is crucial for applications in structural health monitoring and audio forensics.
Purpose of the Study:
- To propose and evaluate a deep learning-based approach for extracting speech signals from vibration response data.
- To compare the performance of fully connected and convolutional neural networks for this task.
- To assess the robustness of the proposed method against various deviations, including position, noise, and boundary conditions.
Main Methods:
- Established a vibroacoustic coupling finite element model with voice signals as the excitation source.
- Utilized vibration acceleration signals from response points for deep learning model training, extracting spectral characteristics.
- Trained and tested fully connected and convolutional neural networks, converting extracted signals back to the time domain using amplitude spectra and phase information.
Main Results:
- Fully connected networks exhibited faster convergence rates and better speech extraction quality compared to convolutional networks.
- The proposed method demonstrated robustness to deviations in vibration response point positions and boundary conditions.
- Speech signal noise significantly impacted extraction quality more than vibration signal noise; performance degraded when both were noisy.
Conclusions:
- The deep learning approach offers a robust and effective solution for speech extraction from vibration signals, surpassing traditional methods in handling deviations.
- Fully connected networks are particularly well-suited for this vibroacoustic system identification task.
- The method shows promise for reliable speech recognition in environments with varying signal quality and system parameters.
More Related Videos
Related Concept Videos
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Echo
Imagine the sound is reflected back to the ears. Assuming that the source is very close to the human, the difference between hearing the two sounds—the emitted sound and the reflected sound—may be more than the minimum time for perceiving distinct sounds. If this is the case,...
Discrete Fourier Transform
Extraction: Advanced Methods
Perception of Sound Waves
The pitch of a sound depends on the frequency and the pressure amplitude of the source. Two sounds of the same...
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...

