揭露目标或背景语音扩展的高频内容的差异性好处
Brian B Monson1, Rohit M Ananthanarayana1, Allison Trine1
1Department of Speech and Hearing Science, University of Illinois Urbana-Champaign, Champaign, Illinois 61820, USA.
The Journal of the Acoustical Society of America
|July 25, 2023
概括
扩展高频 (EHF) 通过提供语音信息或帮助说话者分离来改善语音识别. 目标EHF的可听性有利于语音中的语音场景,但掩盖EHF没有提供隔离优势.
科学领域:
- 听觉神经科学 听觉神经科学
- 语音感知 语音感知
- 信号处理 信号处理
背景情况:
- 已知扩展高频 (EHF;>8 kHz) 有助于语音识别,特别是在复杂的听力环境中,如语音中的语音场景.
- 通过超频频受益于语音感知的确切机制仍然不完全理解,其中可能包括直接的语音信息,加强对低频信息的访问或改进说话者隔离.
研究的目的:
- 调查EHF在语音识别中的好处背后的具体机制.
- 为了确定EHF的好处是否来自于EHF频段内的声学信息,改善对低频声学线索的访问,或说话者分离线索.
主要方法:
- 语音识别测试使用全频段 (FB) 和低通过 (LP;<8 kHz) 语音进行,用于目标和面具说话者.
- 通过独立过目标语音和口罩语音,创建了四个条件,从而产生了EHF匹配和EHF不匹配的场景,其中有一个和两个说话者口罩.
- 一个逐字分析检查了目标词和识别准确度中EHF能量水平之间的关系.
主要成果:
- 语音识别性能在目标语音是全频段 (FB) 和面罩是低通过 (LP) 时是最高的,这种效果在两个说话者面罩时更为明显.
- 当目标语音被低通过 (LP) 时,没有观察到EHF不匹配的好处.
- 逐字分析显示,更高的识别准确度与目标词中增加的EHF能量相关.
结论:
- 目标EHF的可听性对语音识别至关重要,可能通过提供直接的语音信息或增强目标隔离和选择性注意力.
- 在口罩演讲中存在的EHF似乎没有给演讲者带来任何显著的隔离益处.
- 这些发现凸显了在具有挑战性的听觉条件下,为了最佳的语音感知,保护EHF的重要性.
相关概念视频
Aliasing
163
Accurate signal sampling and reconstruction are crucial in various signal-processing applications. A time-domain signal's spectrum can be revealed using its Fourier transform. When this signal is sampled at a specific frequency, it results in multiple scaled replicas of the original spectrum in the frequency domain. The spacing of these replicas is determined by the sampling frequency.
If the sampling frequency is below the Nyquist rate, these replicas overlap, preventing the original...
If the sampling frequency is below the Nyquist rate, these replicas overlap, preventing the original...
163
Masking and Demasking Agents
2.5K
EDTA titrations may necessitate masking and demasking agents to temporarily protect a particular metal ion in a mixture from the EDTA reaction. These agents facilitate the sequential analysis of the metal ions by forming stable complexes with some—but not all—metal ions during certain steps.
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
2.5K
Difference from Background: Limit of Detection
6.4K
The limit of detection (LOD) is the smallest amount of analyte that can be distinguished from the background noise. The LOD value corresponds to the concentration at which the analyte signal is three times larger than the standard deviation of the blank signal. Below this value, the analyte signal cannot be differentiated from the background noise. It is calculated by dividing the calibration slope by 3 times the standard deviation of the blank signals.
The LOD indicates the presence or absence...
The LOD indicates the presence or absence...
6.4K
Upsampling
265
Managing signal sampling rates is essential in digital signal processing to maintain signal integrity. A decimated signal, characterized by a reduced frequency range due to its lower sampling rate, can be upsampled by inserting zeros between each sample. This upsampling process expands the original spectrum and introduces repeated spectral replicas at intervals dictated by the new Nyquist frequency. To refine this zero-inserted sequence, it is passed through a lowpass filter with a cutoff...
265
¹³C NMR: Distortionless Enhancement by Polarization Transfer (DEPT)
1.1K
When proton-coupled carbon-13 spectra are simplified by a broadband proton decoupling technique, structural information about the coupled protons is lost. Distortionless enhancement by polarization transfer (DEPT) is a technique that provides information on the number of hydrogens attached to each carbon in a molecule. While the DEPT experiment utilizes complex pulse sequences, the pulse delay and flip angle are specifically manipulated. The resulting signals have different phases depending on...
1.1K
Interference: Path Lengths
1.3K
Consider two sources of sound, that may or may not be in phase, emitting waves at a single frequency, and consider the frequencies to be the same.
Two special sources may be considered when they are in phase. This can be easily achieved by feeding the two sources from the same source. An example would be synchronizing the two speakers by feeding them with the same source, such as the sound waves produced by a tuning fork. This setup ensures that the two sources have the same frequency and are...
Two special sources may be considered when they are in phase. This can be easily achieved by feeding the two sources from the same source. An example would be synchronizing the two speakers by feeding them with the same source, such as the sound waves produced by a tuning fork. This setup ensures that the two sources have the same frequency and are...
1.3K


