从基于深度学习的振动信号中提取语音
Li Wang1,2,3, Weiguang Zheng1,4, Shande Li1,3
1State Key Laboratory of Digital Manufacturing Equipment and Technology, School of Mechanical Science and Engineering, Huazhong University of Science and Technology, Wuhan, China.
PloS one
|October 25, 2023
概括
本研究介绍了一种深度学习方法,用于从振动信号中提取语音,克服传统的局限性. 完全连接的网络在从振动声学数据的语音识别方面表现出卓越的性能和稳定性.
科学领域:
- 声学和信号处理
- 机器学习和人工智能的人工智能
- 结构动力学和系统识别
背景情况:
- 从振动信号中进行传统的语音提取对模型参数,噪声和边界条件的偏差敏感.
- 现有的方法在从复杂的振动声系统中准确识别语音特征方面面临挑战.
- 开发强大的语音提取方法对于结构健康监测和音频法医学的应用至关重要.
研究的目的:
- 提出和评估基于深度学习的方法,从振动响应数据中提取语音信号.
- 为了比较这个任务的完全连接和卷积神经网络的性能.
- 评估拟议方法对各种偏差的稳定性,包括位置,噪声和边界条件.
主要方法:
- 建立了一个振动声学合有限元模型,用语音信号作为激发源.
- 利用来自响应点的振动加速信号进行深度学习模型训练,提取光谱特征.
- 经过训练和测试的完全连接和卷积神经网络,使用振幅光谱和相位信息将提取的信号转换回时间域.
主要成果:
- 与卷积网络相比,完全连接的网络显示出更快的融合率和更好的语音提取质量.
- 拟议的方法证明了对振动响应点位置和边界条件偏差的稳定性.
- 语音信号噪声明显影响了提取质量,超过了振动信号噪声;当两者都很的时候,性能降低了.
结论:
- 深度学习方法为从振动信号中提取语音提供了强大而有效的解决方案,超越了处理偏差的传统方法.
- 完全连接的网络特别适合于这种振动声系统识别任务.
- 该方法在信号质量和系统参数变化的环境中显示出可靠的语音识别的前景.
更多相关视频
相关概念视频
Classification of Signals
484
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
484
Echo
515
The human ear cannot distinguish between two sources of sound if they happen to reach within a specific time interval, typically 0.1 seconds apart. More than this, and they are perceived as separate sources.
Imagine the sound is reflected back to the ears. Assuming that the source is very close to the human, the difference between hearing the two sounds—the emitted sound and the reflected sound—may be more than the minimum time for perceiving distinct sounds. If this is the case,...
Imagine the sound is reflected back to the ears. Assuming that the source is very close to the human, the difference between hearing the two sounds—the emitted sound and the reflected sound—may be more than the minimum time for perceiving distinct sounds. If this is the case,...
515
Discrete Fourier Transform
305
The Discrete Fourier Transform (DFT) is a fundamental tool in signal processing, extending the discrete-time Fourier transform by evaluating discrete signals at uniformly spaced frequency intervals. This transformation converts a finite sequence of time-domain samples into frequency components, each representing complex sinusoids ordered by frequency. The DFT translates these sequences into the frequency domain, effectively indicating the magnitude and phase of each frequency component present...
305
Extraction: Advanced Methods
463
Metal ions can be separated from one another by complexation with organic ligands–the chelating agent– to form uncharged chelates. Here, the chelating agent must contain hydrophobic groups and behave as a weak acid, losing a proton to bind with the metal. Since most organic ligands used in this process are insoluble or undergo oxidation in the aqueous phase, the chelating agent is initially added to the organic phase and extracted into the aqueous phase. The metal-ligand complex is...
463
Perception of Sound Waves
4.5K
The human ear is not equally sensitive to all frequencies in the audible range. It may perceive sound waves with the same pressure but different frequencies as having different loudness. Moreover, the perception of sound waves depends on the health of an individual's ears, which decays with age. The health of one's ears may also be affected by regular exposure to loud noises.
The pitch of a sound depends on the frequency and the pressure amplitude of the source. Two sounds of the same...
The pitch of a sound depends on the frequency and the pressure amplitude of the source. Two sounds of the same...
4.5K
Perceiving Loudness, Pitch, and Location
220
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
220


